Skip to content

feat(verification): committed gates, protocol triage, and six command entry points - #64

Draft
ManSio wants to merge 5 commits into
mainfrom
feat/protocol-triage-and-t10-guards
Draft

ManSio wants to merge 5 commits into
mainfrom
feat/protocol-triage-and-t10-guards

Conversation

@ManSio

@ManSio ManSio commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Brings the verification gates and the six protocol commands into the repository so they version with the code they audit. Both previously lived under ~/.config, outside any git repository, which made them unusable in CI and for anyone else.

MSCodeBase Agent added 5 commits September 30, 2026 01:35
…pt defaults

F0c tail: .local/ was gitignored but still tracked - rm --cached (disk files kept, CI unaffected). reconstruct_judge_cot DEFAULT_DB via Path.home() (same resolution on owner box, portable elsewhere); WORKDIR kept as historical DB filter. o1/p2 holdout gate usage examples use <repo-root> placeholder. R1 rotation: KNOWN_ISSUES.md 448->197 lines (33 closed blocks deduped against archive, bodies verified present). Diary synced same commit.
…mber

Three rate tools returned or crashed on an empty population:

- scripts/e2e_quality_search.py: hit@1/hit@5 divided by len(rows) with no
  guard; an empty result set raised ZeroDivisionError. Now exits 2 with a
  diagnosable message. avg_ms in the same function was already guarded.
- experiments/context_engine/compose_eval.py: wrong_ratio returned 0.0 on an
  empty set. The division was guarded, but 0.0 is the PERFECT score, so this
  produced a plausible false PASS. Now raises.
- experiments/misc_probes/exp_vacuous_scan.py: REPO_ROOT resolved relative to
  the script instead of the package, so the scanned path did not exist and the
  scan reported 0 total. Now resolves correctly and exits 2 on an empty total.

Adds scripts/triage_protocol_findings.py, which measures the false-positive
share of the protocol guard: 5 of 8 findings (62.5%) were noise, leaving 2 true
positives and 1 partial. Two automated triage attempts are recorded as rejected:
a keyword scan missed guarded files, and a division regex cited filesystem
paths as rate sites. They erred in opposite directions, so neither is usable.

Tests: tests/test_t10_empty_population.py holds a negative control (empty input
must be refused) and a positive control (non-empty must still compute).
… the code

The verification gates lived in the agent's personal config directory, outside any
git repository. That makes them useless for CI and for anyone else: a guard that
is not committed is not a guard. They now live in tools/verification/ and travel
with the code they audit.

Paths are derived from __file__ rather than hardcoded, so a clone at any location
runs the same suite. heldout_relocation.py enforces that: it greps for
author-absolute paths and runs every gate from an unrelated working directory.

heldout_g5.py was rewritten after three broken versions. Each mutated state that
outlived the test — a corrupted manifest restored from bytes, a "pristine" baseline
snapshotted from an already-dirty file, and a gate sabotaged in place and then
restored with the sabotaged text, which left a 0-byte gate. All cases now run in
their own temp directory and the real manifest is never written to. It also
sabotages the gate on purpose and requires the suite to notice, so the harness
itself is provably falsifiable.

G5 currently reports 0.00% coverage: 0 of 753 candidates classified. That is the
honest reading, not a broken gate.
tests/test_planted_break_gate.py writes
experiments/planted_break/results.json on every run, so a full pytest pass left
the working tree dirty and the change was a timestamp nobody had reviewed. It is
written by the test and never read back as a fixture, so there is nothing to lose
by untracking it.

.gitignore now carries the path, with a note on why, so the next person does not
re-add it and wonder why git status is never clean.
The command entry points lived in ~/.config/opencode/commands, outside any git
repository. A command file that is not committed is lost with the machine,
overwritten by a reinstall, and unreadable by anyone else — and these commands
invoke the gates in tools/verification, so a command and its guards must share a
history. They now live in .opencode/command/ next to the existing tracked plugin.

Each command is a procedure that refuses one specific false conclusion: a summary
mistaken for a source, a number with no control, a denominator authored so that
coverage returns 100% by construction, a ship decision justified by green tests, a
fix with no guard, and a ✅ with no statement of where it was verified.

heldout_cli_contract.py pins the exact invocations those command files cite. The
commands name real script paths, so a change to the CLI shape would silently turn
them into prose that cannot be run; the check makes that a failing step in
run_all.py rather than a discovery made later.
@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant