Skip to content

docs: dogfooding runbook, and correct the CI repair gate's spend figures (U7) - #26

Merged
Steel-tech merged 1 commit into
mainfrom
docs/dogfooding-runbook
Sep 28, 2026
Merged

Steel-tech merged 1 commit into
mainfrom
docs/dogfooding-runbook

Conversation

@Steel-tech

Copy link
Copy Markdown
Contributor

Implements U7 of the factory-quality follow-ups plan (PR #21): R14, KTD1's revisit thresholds, the KTD5 wall-clock note, and KTD8.

What the runbook covers (docs/dogfooding.md)

  • 0. What the build must contain. jig report (feat: quality report — GET /api/report and jig report (U2, U3) #24), spend on every exit (feat(engine): record agent spend on every phase exit (U1) #25), publish history (feat(worker): keep publish history across publish-only retries (U8) #22), and optionally re-runs (U4). Each row says what the report gets wrong without it.
  • 1. Preflight.
    • A one-send smoke.yaml run proves runtime auth in an ephemeral HOME, including the macOS keychain case.
    • gh auth status must show repo and workflow; the fine-grained token equivalents are listed too.
    • The real log read (--allow-escape-sequences) and a real per-job POST …/actions/jobs/{id}/rerun.
    • CI is green on main.
    • A review of the target's secrets, variables, and workflows for unattended exposure. It covers the env-block leak and values derived from secrets that GitHub does not mask.
  • 2. Trial protocol.
    • A dedicated serve --data / worker --data pair, so the report counts only trial jobs.
    • Job selection, with a minimum of 20 CI-waiting attempts before any rate is used.
    • Sizing against the attempt ceiling. (1+K)·T_chain + (1+K+R)·T_ci ≤ 4h, because the ceiling is fixed at protocol.MaxAttemptDuration: operators can't raise it, only size inside it.
    • A spend-ceiling check with jig report --json before each batch. It charges unmetered sends at the roster's max budget_usd.
    • A per-job record from GET /api/jobs/{id} and agent_end events, plus the operator's own verdict on each PR.
  • 3. Results. An exit table whose rows map to real jig report --json fields, with a jq one-liner. The decision thresholds:
    • KTD1's ci_timeout_rate and retry_still_red_rate above 0.10;
    • a flaky-pass rate above 0.10 that suggests masking;
    • a recurring flaky check;
    • person-fixed accepts outnumbering repaired ones;
    • clean accepts the operator had to rewrite.
  • 4. Lessons from the CI repair exit gate.

Merge order

This documents jig report from PR #24 and publish.ci.rerun from the U4 PR, so it should merge after both. U4 had no branch when this was written, so re-runs are described from the plan (R5–R7, KTD5), and the doc points at the README for the syntax as shipped. If U4 changes the field name or the budget range, update section 2.1.

Correction

The CI repair plan's Exit gate section now says that its spend figures ($3.03/$5.03/$1.89, $9.95, $13.08) are upper bounds. U1 (PR #25) found that Claude Code's total_cost_usd is cumulative per session, and v0.2.0 summed it per send, so resumed sessions were overcounted. The figures themselves are unchanged.

Verification

  • just check passes.
  • Field names were checked against internal/controlplane/report.go on feat/report-api, and the flags against cmd/jig/report.go and def.go.
  • The per-job and table jq programs were run against sample JSON. job_spend was run under bash and zsh against a stubbed curl.
  • The gh run list / gh run view / log-read commands were run against this repo. The re-run POST was not run, since it re-runs a real job.

🤖 Generated with Claude Code

…res (U7)

docs/dogfooding.md takes an operator from preflight to a results table
for a factory.yaml trial: runtime auth proven by a one-send smoke run,
gh scopes and the real log-read and job re-run calls, green CI on main,
and the target's Actions secrets and variables reviewed for unattended
exposure. The protocol sizes the chain against the fixed 4h attempt
ceiling, checks the spend ceiling with jig report before each batch, and
records per-job outcomes, including the PR review jig cannot see. The
exit table maps each row to a jig report --json field, and the decision
thresholds carry KTD1's revisit rates and a flaky-pass masking rate.

The CI repair plan's exit gate spend figures are upper bounds: Claude
Code's total_cost_usd is cumulative per session (PR #25), and v0.2.0
summed it per send.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 27 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: d780c0fb-7efd-4ff3-95d6-528145959544

📥 Commits

Reviewing files that changed from the base of the PR and between 8c542bb and ee49650.

📒 Files selected for processing (3)
  • README.md
  • docs/dogfooding.md
  • docs/plans/2026-09-27-001-feat-ci-repair-plan.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Steel-tech
Steel-tech merged commit a4ade80 into main Sep 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant