Conversation
added 5 commits
September 30, 2026 01:35
…pt defaults F0c tail: .local/ was gitignored but still tracked - rm --cached (disk files kept, CI unaffected). reconstruct_judge_cot DEFAULT_DB via Path.home() (same resolution on owner box, portable elsewhere); WORKDIR kept as historical DB filter. o1/p2 holdout gate usage examples use <repo-root> placeholder. R1 rotation: KNOWN_ISSUES.md 448->197 lines (33 closed blocks deduped against archive, bodies verified present). Diary synced same commit.
…mber Three rate tools returned or crashed on an empty population: - scripts/e2e_quality_search.py: hit@1/hit@5 divided by len(rows) with no guard; an empty result set raised ZeroDivisionError. Now exits 2 with a diagnosable message. avg_ms in the same function was already guarded. - experiments/context_engine/compose_eval.py: wrong_ratio returned 0.0 on an empty set. The division was guarded, but 0.0 is the PERFECT score, so this produced a plausible false PASS. Now raises. - experiments/misc_probes/exp_vacuous_scan.py: REPO_ROOT resolved relative to the script instead of the package, so the scanned path did not exist and the scan reported 0 total. Now resolves correctly and exits 2 on an empty total. Adds scripts/triage_protocol_findings.py, which measures the false-positive share of the protocol guard: 5 of 8 findings (62.5%) were noise, leaving 2 true positives and 1 partial. Two automated triage attempts are recorded as rejected: a keyword scan missed guarded files, and a division regex cited filesystem paths as rate sites. They erred in opposite directions, so neither is usable. Tests: tests/test_t10_empty_population.py holds a negative control (empty input must be refused) and a positive control (non-empty must still compute).
… the code The verification gates lived in the agent's personal config directory, outside any git repository. That makes them useless for CI and for anyone else: a guard that is not committed is not a guard. They now live in tools/verification/ and travel with the code they audit. Paths are derived from __file__ rather than hardcoded, so a clone at any location runs the same suite. heldout_relocation.py enforces that: it greps for author-absolute paths and runs every gate from an unrelated working directory. heldout_g5.py was rewritten after three broken versions. Each mutated state that outlived the test — a corrupted manifest restored from bytes, a "pristine" baseline snapshotted from an already-dirty file, and a gate sabotaged in place and then restored with the sabotaged text, which left a 0-byte gate. All cases now run in their own temp directory and the real manifest is never written to. It also sabotages the gate on purpose and requires the suite to notice, so the harness itself is provably falsifiable. G5 currently reports 0.00% coverage: 0 of 753 candidates classified. That is the honest reading, not a broken gate.
tests/test_planted_break_gate.py writes experiments/planted_break/results.json on every run, so a full pytest pass left the working tree dirty and the change was a timestamp nobody had reviewed. It is written by the test and never read back as a fixture, so there is nothing to lose by untracking it. .gitignore now carries the path, with a note on why, so the next person does not re-add it and wonder why git status is never clean.
The command entry points lived in ~/.config/opencode/commands, outside any git repository. A command file that is not committed is lost with the machine, overwritten by a reinstall, and unreadable by anyone else — and these commands invoke the gates in tools/verification, so a command and its guards must share a history. They now live in .opencode/command/ next to the existing tracked plugin. Each command is a procedure that refuses one specific false conclusion: a summary mistaken for a source, a number with no control, a denominator authored so that coverage returns 100% by construction, a ship decision justified by green tests, a fix with no guard, and a ✅ with no statement of where it was verified. heldout_cli_contract.py pins the exact invocations those command files cite. The commands name real script paths, so a change to the CLI shape would silently turn them into prose that cannot be run; the check makes that a failing step in run_all.py rather than a discovery made later.
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings the verification gates and the six protocol commands into the repository so they version with the code they audit. Both previously lived under ~/.config, outside any git repository, which made them unusable in CI and for anyone else.