docs: dogfooding runbook, and correct the CI repair gate's spend figures (U7) - #26
Merged
Merged
Conversation
…res (U7) docs/dogfooding.md takes an operator from preflight to a results table for a factory.yaml trial: runtime auth proven by a one-send smoke run, gh scopes and the real log-read and job re-run calls, green CI on main, and the target's Actions secrets and variables reviewed for unattended exposure. The protocol sizes the chain against the fixed 4h attempt ceiling, checks the spend ceiling with jig report before each batch, and records per-job outcomes, including the PR review jig cannot see. The exit table maps each row to a jig report --json field, and the decision thresholds carry KTD1's revisit rates and a flaky-pass masking rate. The CI repair plan's exit gate spend figures are upper bounds: Claude Code's total_cost_usd is cumulative per session (PR #25), and v0.2.0 summed it per send. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Warning Review limit reachedNext included review available in 27 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements U7 of the factory-quality follow-ups plan (PR #21): R14, KTD1's revisit thresholds, the KTD5 wall-clock note, and KTD8.
What the runbook covers (
docs/dogfooding.md)jig report(feat: quality report — GET /api/report and jig report (U2, U3) #24), spend on every exit (feat(engine): record agent spend on every phase exit (U1) #25), publish history (feat(worker): keep publish history across publish-only retries (U8) #22), and optionally re-runs (U4). Each row says what the report gets wrong without it.smoke.yamlrun proves runtime auth in an ephemeral HOME, including the macOS keychain case.gh auth statusmust showrepoandworkflow; the fine-grained token equivalents are listed too.--allow-escape-sequences) and a real per-jobPOST …/actions/jobs/{id}/rerun.main.serve --data/worker --datapair, so the report counts only trial jobs.(1+K)·T_chain + (1+K+R)·T_ci ≤ 4h, because the ceiling is fixed atprotocol.MaxAttemptDuration: operators can't raise it, only size inside it.jig report --jsonbefore each batch. It charges unmetered sends at the roster's maxbudget_usd.GET /api/jobs/{id}andagent_endevents, plus the operator's own verdict on each PR.jig report --jsonfields, with ajqone-liner. The decision thresholds:ci_timeout_rateandretry_still_red_rateabove 0.10;Merge order
This documents
jig reportfrom PR #24 andpublish.ci.rerunfrom the U4 PR, so it should merge after both. U4 had no branch when this was written, so re-runs are described from the plan (R5–R7, KTD5), and the doc points at the README for the syntax as shipped. If U4 changes the field name or the budget range, update section 2.1.Correction
The CI repair plan's Exit gate section now says that its spend figures ($3.03/$5.03/$1.89, $9.95, $13.08) are upper bounds. U1 (PR #25) found that Claude Code's
total_cost_usdis cumulative per session, and v0.2.0 summed it per send, so resumed sessions were overcounted. The figures themselves are unchanged.Verification
just checkpasses.internal/controlplane/report.goonfeat/report-api, and the flags againstcmd/jig/report.goanddef.go.jqprograms were run against sample JSON.job_spendwas run under bash and zsh against a stubbedcurl.gh run list/gh run view/ log-read commands were run against this repo. The re-run POST was not run, since it re-runs a real job.🤖 Generated with Claude Code