chore(release): 0.21.0 — ADR-065 envelope restoration + typed /gate refusals - #115
Merged
Merged
Conversation
BUDGET_WORKFLOW_BLOCKED and BUDGET_CACHE_EXCEEDED are registered in the backend's GateErrorCode::all() (error_codes.rs:666-667, both 402) and the backend logged BUDGET_WORKFLOW_BLOCKED x389 in production before they were registered there at all. The SDK catalog was never updated to match, so the typed dispatcher could not classify either: "BUDGET_WORKFLOW_BLOCKED" in _V3_ERROR_CODE_MAP -> False The code therefore fell through to the base-class drift tier, which preserves the literal string for visibility but yields no typed budget class. A caller branching on NullRunBudgetError to mean "stop spending" saw an untyped block instead. This drift is invisible in practice because the /gate pre-flight used to raise NullRunBudgetError unconditionally for every refusal, so nothing downstream ever consulted the catalog. That hardcode is removed in the companion commit, which is what exposes the gap. Found by QA cycle RUN_ID 20261002T0826 (TC-4), against production build e8d811a0. Tests: 1603 passed, 1 skipped. ruff clean. mypy clean for the files touched here (the 6 remaining errors are in instrumentation/auto.py and reproduce identically without these changes).
check_workflow_budget raised NullRunBudgetError for EVERY block. A
rate-limit refusal, a policy tool-block and a workflow-inactive all
reached the caller as NR-B004 "budget exhausted". The typed
dispatcher already existed and was wired into Runtime.execute; the
/gate pre-flight simply never called it, so the two paths could not
disagree about what a refusal means.
Observed on production (QA cycle RUN_ID 20261002T0826, TC-4,
build e8d811a0) against a workflow with rate_limit_per_minute=5:
[0..4] ALLOW
[5] NullRunBudgetError code=NR-B004
"blocked: RATE_LIMIT_EXCEEDED"
The gate was correct — 5 allows then a block at index 5. The
classification was not: an operator reading NR-B004 goes to raise a
spend cap when the actual cause is a throttle policy, and a caller
catching NullRunBudgetError to mean "stop spending" stops for the
wrong reason.
Four changes, each independently necessary:
1. The pre-flight routes through _build_block_exception.
2. The dispatcher resolves the wire code the way the BACKEND resolves
the HTTP status for the same body: an ordered candidate list —
details["error_code"], then the top-level error_code that
Transport.check already copies onto its 4xx dict, then
explanation (which the Block dispatcher binds reason_code into).
Matching the backend's order is what keeps the status and the
exception class from disagreeing about which code a refusal is.
An explicitly-sent code still reaches the base-class drift tier
when unregistered, preserving test_unknown_wire_code_falls_back_
to_base; only explanation-as-candidate is gated on registration,
so English prose never presents as a wire code.
3. The dispatcher handles all three catalog families, not just the
NullRunBlockedException one. The transport family
(RateLimitError, NullRunBackendError, NullRunExecutionNotFoundError)
takes (message, source, endpoint, ...), infra-not-transport
(NullRunRateLimitRedisError) takes the bare message. Passing the
decision-shaped kwargs raised TypeError and let the refusal escape
untyped. Dispatch is by signature, so a future catalog entry
cannot silently pick the wrong shape. source is never forwarded
through **details — it collides with the keyword the class passes
down itself.
4. cost_limit_exceeded is bumped only for NullRunBudgetError. It is
the operator's "budget cap is biting" counter; bumping it for a
rate-limit block is the second half of the same mislabelling.
The pre-flight keeps its own caller contract on the legacy tier: a
response with no machine-readable code still raises
NullRunBudgetError, pinned by test_real_block_still_honored,
test_enforcement_4xx_still_raises_budget_error and
test_block_response_does_not_infect_subsequent_track. The shared
dispatcher stays on the base class for a type guessed from English
(test_legacy_keyword_path_budget) — a subclass chosen by
substring-matching claims more confidence than the wire gave. The two
callers want different things from the same dispatcher output, so
the widening happens at the caller that promises it.
Two existing tests changed, both asserting the artefact of the
hardcode rather than a behaviour:
- test_403_breaker_trip_stops_the_agent / test_403_is_not_retried
expected NullRunBudgetError for CIRCUIT_BREAKER_TRIPPED, a
category-halt refusal that is absent from the catalog and now
correctly surfaces as the base class with the wire code preserved.
Their actual subject — honoured, not retried, reason text intact —
is unchanged and still asserted.
- test_legacy_keyword_path_budget is untouched.
New: tests/test_gate_block_typed_dispatch.py, 6 cases covering the
typed rate-limit classes, the tool-block-is-not-budget half, the
budget path, and the legacy tier. Verified failing on the old code
first (4 failed, 2 passed) before any change.
Tests: 1603 passed, 1 skipped. ruff clean. mypy clean for the files
touched (the 6 remaining errors are in instrumentation/auto.py and
reproduce identically without these changes).
DEF-TC6-005 (QA RUN_ID 20261002T0826, SDK 0.20.0). TC-12
`approval_granted` against production reported:
STATUS_OK=NullRunStatus(..., ws_connected=None, ...)
with the listener thread alive and a real connection present 0.3s
after init():
t=0.3s conn present type=WebSocketConnection
has is_open: False
attrs: ['url', 'headers', 'api_key', 'secret_key',
'on_state_change', 'on_policy_invalidated',
'on_key_rotated', 'on_approval_resolved', '_conn',
'_running', '_receive_task', '_reconnect_task',
'_closed', '_consecutive_reconnect_failures',
'_last_version']
The same run logged a full WS lifecycle (CLOSE 1000 / EOF / "WebSocket
connection closed"), so the channel did come up.
`status()` computed the field as
ws_connected = getattr(self._ws_connection, "is_open", None)
`WebSocketConnection` never had an `is_open` attribute — its liveness
flag is `_running`, set True in `_connect` (transport_websocket.py:243)
and cleared by the receive loop's `finally` (:283). The `getattr`
default therefore fired on every call and the field was structurally
pinned to `None`. `is_open` appears exactly once in the SDK: on the
reading side, with no writer, no test and no producer.
This matters because the three states are distinct and the collapse
hides the one that matters — a listener that started and then
dropped. `None` means "never established", so a channel that came up
and died was indistinguishable from one that never existed, and
`status()` — the SDK's only introspection surface — could not report
a working push channel at all.
Read `_running`. The `getattr` default stays so a connection object
without the flag degrades to `None` rather than raising, and all
three states still report distinctly: never-established `None`,
live `True`, dropped-or-shutdown `False`.
Tests: tests/test_status_ws_connected_reads_real_attribute.py, four
cases. The fourth is the one that would have caught this: it asserts
against a really-constructed `WebSocketConnection` that the attribute
`status()` reads exists, and that `is_open` does not — so a future
rename of `_running` fails here rather than silently returning `None`
again in production. A stubbed connection cannot catch it, because
the stub would carry whatever attribute the test author assumed.
Verified by mutation: reverting the read to `is_open` reds the two
state tests.
pytest: 1607 passed, 1 skipped, 0 failed.
ruff + mypy: clean.
DEF-TC6-006 (QA RUN_ID 20261002T0826, 2026-10-02). TC-15
consume_overbudget against production observed a 422
CONSUME_OVERBUDGET on the wire, "event dropped" in the log, and
TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0}
returned to the probe.
_route_track wrapped transport.track_single in a bare
`except Exception` that logged at WARNING and returned. The
transport layer had already classified the response — a 422
CONSUME_OVERBUDGET body becomes a typed
NullRunConsumeOverbudgetError carrying reserved_cents,
actual_cost_cents, max_allowed_cents and epsilon_cents
(transport.py:2939) — and the catch discarded it.
Two things make the bare catch wrong rather than blunt:
* The drop-and-log policy it implements is the one the ADR-008
table states for the `/track batch path (legacy)`, a NETWORK
error. The v3 single path has no such row.
* The same docstring says the SDK "does NOT silently fail-OPEN
on a wire 4xx/5xx that names an enforcement failure", naming
/track among the handlers that raise. A refused consume names
one.
Re-raise NullRunDecision after the existing cache invalidation and
telemetry, so the blast-radius mitigation (DEF-CACHE-STALE-ALLOW-
AFTER-OVERBUDGET) is not traded away for the reporting fix. Scope
is NullRunDecision, not Exception: connection errors and 5xx name
no enforcement failure and keep dropping, because raising there
would freeze the agent loop on a dead backend — the exact failure
the fail-OPEN rows exist to prevent.
ADR-008 table gains a row for the v3 single path; README fail-open
claim unchanged (this narrows a swallow, it does not widen
enforcement).
tests/test_route_track_enforcement_propagates.py (new, 5 tests):
3 fail on the pre-fix code, 2 pin the drop-and-log half so the fix
cannot over-reach. The two existing cache-invalidation tests are
strengthened, not relaxed — they now assert the raise AND the
invalidation.
Full suite: 1612 passed, 1 skipped, 0 failed. ruff + mypy clean.
`check_workflow_budget` has computed the MCP forwarding fields since
the MCP integration landed — it reads `get_call_mcp_class()` and
`get_call_mcp_annotations()` off the call context and sets them on
`check_req` (`runtime.py:2383-2388`) — but `Transport.check` never
sent `check_req`. It rebuilds the `/gate` body from an explicit
allowlist (`transport.py:1524-1542` plus the conditional forwards
that follow) and neither key was on it, so both values were dropped
without a word.
Observed on the wire against prod `e8d811a0` under QA RUN_ID
20261002T0826 (TC-29). With
set_mcp_tool_context(tool_class="mcp",
annotations={"read_only": False,
"destructive": True,
"open_world": False})
followed by one `check_workflow_budget()`, the captured `/gate` body
was:
['action_digest', 'check_type', 'estimated_tokens', 'execution_id',
'idempotency_key', 'input', 'mode', 'model', 'operation_id',
'organization_id', 'stream', 'tool', 'tools', 'trace_id']
— neither field present. The public `set_mcp_tool_context` API and
the `toolbox.mcp` auto-classification path were dead end to end: the
SDK could never tell the gate a tool was `destructive` or `read_only`.
The backend is complete and states the contract it expects
(`gate/internal.rs:318-341`): "SDKs that recognise an MCP server
cache the `tools/list` entry and pass the canonical `Mcp` class plus
the corresponding `McpAnnotations`". `effective_tool_class()` falls
back to `classify_tool` on the raw tool string when the field is
absent, so this is a dead feature with a misleading API rather than
a live bypass today — but on the day the server-side flag flips,
destructive MCP tools will silently degrade to name-based
classification.
Guarded on `is not None`, not on key presence. The backend pins the
negative case too (`internal.rs:8291-8295` asserts `tool_class=None`
and `mcp_annotations=None` must not appear in the JSON), and an
absent annotation means "unknown", not "false"
(`internal.rs:334-339`). Serialising `null` would be a different
value carrying a different meaning.
Tests assert the captured POST body rather than pinning source text,
so a future refactor that reintroduces the drop at a different line
still fails. Verified failing on the pre-fix code (3 failed, 2
passed) and passing after (5 passed); full suite 1617 passed, 1
skipped, 0 failed.
Not pushed.
DEF-TC21-001, found 2026-10-02 under QA RUN_ID 20261002T0826 (TC-21).
Three shipped places told users a property the class has not had since
0.16.6:
* `docs/errors/NR-W002.md` — "Exception class: `WorkflowKilledInterrupt`
(subclass of `BaseException`, NOT `Exception`)" and "`except Exception`
will NOT catch this signal by design". Also pointed at
`docs/kill-contract.md` §6, a file that does not exist.
* `src/nullrun/breaker/exceptions.py` — the `NullRunError` docstring
repeated the claim verbatim.
* the same module's class catalog listed `WorkflowKilledException` as a
live public class; it was deleted in the 0.18.5 sweep.
Verified at runtime on 0.20.0:
WorkflowKilledInterrupt -> [WorkflowKilledInterrupt, NullRunError,
BreakerError, Exception, BaseException, object]
`BreakerError` derives from `Exception` (`exceptions.py:5`), so a broad
`except Exception:` in host code swallows a kill — the opposite of what
the docs promised, and the failure mode is silent: the workflow looks
alive and the agent keeps calling tools.
The behaviour itself is correct and deliberate. `9877c34` (0.16.6)
reparented the class so agent recovery can catch a kill and surface the
structured `error_code` / `user_action`; `tests/test_decision_split.py`
documents the override. Only the prose was wrong, and it was wrong in the
direction that would lead the next maintainer to revert working code.
`tests/test_exception_hierarchy.py` had the same disease: the test was
named `test_killed_interrupt_does_not_inherit_from_exception` while
asserting `issubclass(WorkflowKilledInterrupt, Exception)`. The
assertion was right; the name said its opposite. Renamed.
Full suite 1617 passed, 1 skipped, 0 failed.
Not pushed.
ADR-065 decision step 1 + verification step 1. Since 0.18.5 every @Protect call sent a constant {"kind":"none"} sentinel. A constant hashes to a constant, so the digest stored on the approval row at /gate and the one the backend recomputes at /execute always agreed - while binding the approval to nothing at all, which is the weakness NR-010 was raised to close. The result is that an operator approves an action and the agent still cannot run it (DEF-TC14-002). Adds ToolCallParams + BusinessImpact.tool_call(), emitting exactly what the backend's internally-tagged serde produces. The extractor_id/extractor_version fields have no skip_serializing_if on the Rust side, so they are inside the hashed bytes and must be present here too. compute_action_digest is refactored to delegate to a new canonical_bytes() so the canonical form is assertable directly. A digest is opaque - it shows that two payloads disagree, never how. The first version of the test asserted on a COPY of the serializer inside the test file, and an ensure_ascii=True mutation left it green; the mutation pass caught that and the assertions now run against the production path. Also restores a pin that aee8110 (0.18.5 deprecation sweep) deleted along with the money/tool_call constructors. That commit removed tests/test_business_impact.py while leaving the module docstring claiming "The digest is pinned by tests/test_business_impact.py", the same docstring/code disagreement that made DEF-TC14-002 read as a backend bug rather than a missing cross-repo contract. The shared fixture digest is byte-identical to the backend's DIGEST_FIXTURE_HEX_TOOL_CALL (9975a8b75a436fb78b9d141b9e0c0a90838c1243d78119b304ae6ed0526966a6). Mutation-verified: ensure_ascii=True, EXTRACTOR_VERSION drift and a kind-string typo each fail the tests that own them. pytest -q: 1656 passed, 1 skipped, 0 failed.
ADR-065 decision step 2. /gate and /execute are two HTTP calls made from two different places at two different times, and the backend compares the digest it RECOMPUTES at /execute against the one it STORED at /gate (payload_binding.rs:163, orchestrator.rs:1511). They therefore have to read the same envelope. A contextvar is the right home for it: set_call_context and set_mcp_tool_context already use this exact pattern for the model, the tool list and the MCP annotations, and it keeps the envelope out of the signature of every function in between. None is a distinct third state from an envelope that says no impact. Collapsing the two is what let a constant sentinel stand in for a tool call for a full release cycle. Steps 3 (/gate reads this instead of minting a sentinel) and 4 (@Protect builds it) are not in this commit. Mutation-verified: a getter that returns a fresh no_impact(), a setter that discards its Token, and a reset that does not restore each fail the tests that own them. pytest -q: 1665 passed, 1 skipped, 0 failed.
ADR-065 decision step 3. check_workflow_budget unconditionally built no_impact() and hashed that. Since the envelope was a constant, the digest the backend stored on the approval row was the same for every tool call, so an approval granted for one tool was equally 'valid' for any other -- the binding the digest exists to provide was not there. The pre-flight now reads the envelope the context already holds and sends it alongside the digest. Transport.check is an allowlist BUILDER, not a pass-through (DEF-TC29-001), so business_impact had to be added to it explicitly or it would have been dropped in transit exactly as approval_id and tool_class were. The backend already declares the field on GateRequest (internal.rs:216) and round-trips it through serde (internal.rs:6947). A context with no envelope is an LLM check with no tool to name, for which no_impact() is the correct and honest answer -- it is not a stand-in for a tool call, and the backend tells the two apart. The tests assert the CAPTURED POST BODY and include a class that goes through check_workflow_budget rather than calling Transport.check with a hand-built envelope. The first version of this file only did the latter, and a mutation restoring the sentinel in runtime.py passed all six of its tests. Mutation-verified: restoring the sentinel, and dropping business_impact from the builder, each fail the tests that own them. The sentinel mutation's own failure output is the defect -- both /tools hash 0049d93a36f0710269a6deb733ca78d57a770ef640a2698d0fddaa9653b7c3de. pytest -q: 1675 passed, 1 skipped, 0 failed.
ADR-065 decision step 4. This is the change that makes the post-approval re-entry reachable at all. @Protect built its envelope inside _run_tool_policy_gate, which runs as step 4 -- after the step-2 /gate pre-flight. The /gate body therefore carried a no_impact() sentinel while the /execute body carried a tool_call envelope, and the backend stores the digest from the FIRST and recomputes it from the SECOND. The envelope is now built before step 2 and both calls read it off the call context. params is the masked argument bag keyed as the /execute input already keys it ({"args": [...], "kwargs": {...}}), so the digest covers exactly what the operator's approval card showed. Two reasons for that shape rather than kwargs-only: a tool called purely positionally would otherwise have params={} and charge_card("x", 50) would be the same action as charge_card("x", 5000); and a digest over unmasked arguments would bind the grant to values nobody approved. A build failure falls back to no_impact() and logs. That is fail-OPEN on metadata and deliberately so: the backend then stores a digest that binds nothing and refuses the re-entry, which is the correct posture, whereas raising out of the decorator would take down a tool call over a metadata field. Also corrects the docstring at decorators.py:869-875 that claimed business_impact was "now always {"kind": "none"}" -- the code and the documentation disagreed, and the code is what shipped. That disagreement is what made DEF-TC14-002 read as a backend bug. The tests capture BOTH request bodies and recompute the digest the way the server does, from an independent implementation rather than by calling compute_action_digest on both sides. Mutation-verified, each independently: /execute falling back to no_impact (9 fail), the envelope omitted from /execute (9 fail), params dropping positional args (2 fail), and -- the one that first appeared to pass -- the envelope built after the pre-flight instead of before (6 fail). The first attempt at that last mutation added a second set_call_impact rather than moving the first, so it changed nothing and the suite stayed green; the reordered version is what pins the ordering ADR-065 calls out as load-bearing. pytest -q: 1690 passed, 1 skipped, 0 failed.
ADR-065 retires the none sentinel FROM THE TOOL PATH, which is only true modulo _build_call_impact's fallback: a tool name the backend's validator rejects (non-ASCII, over 128 bytes) degrades to no_impact() and logs. The tool body still runs; the approval carries no trust binding and the server refuses the re-entry. That branch was previously untested, which made "retired from the tool path" a claim rather than a fact. A realistic trigger is a non-ASCII tool name -- the backend requires printable ASCII (business_impact.rs:320-323), so such a tool cannot be described by a valid envelope at all. Mutation-verified: replacing the fallback with a fabricated tool_call envelope fails 3 of the 4 new tests. pytest -q: 1694 passed, 1 skipped, 0 failed.
…efusals
Cuts `v0.21.0` — the SDK-side closure of `DEF-TC14-002` (QA cycle
RUN_ID `20261002T0826`, 2026-10-02). Two QA cycles in a row found
the same defect from two angles: 0.20.0's audit said "operator
approves an action and the agent still cannot run it", and this
cycle's TC-4 said "/gate blocks on a rate limit but the SDK raises
`NullRunBudgetError` so callers branch on the wrong cause". Both
were the same root: the SDK was sending a constant sentinel instead
of the business impact envelope the backend was looking for, and
was classifying refusals by what the SDK's wrapper assumed rather
than by what the wire said.
The load-bearing fix is the five-step restoration of the
`BusinessImpact` envelope (ADR-065). Since 0.18.5 every `@protect`
call sent a constant `{"kind":"none"}` sentinel; a constant hashes
to a constant, so the digest the backend stored on the approval row
at `/gate` matched the digest it recomputed at `/execute` for every
tool — while binding the approval to nothing at all. The operator's
approval card did not correspond to any particular action, and the
backend's refuse-the-reentry check had no data to refuse *against*.
## Migration (read first if upgrading from 0.20.x)
Three things differ from 0.20.0. None are silent on a healthy
configuration, but each can be a working loop becoming a throwing
one for code that suppressed the failure before.
1. `BusinessImpact.tool_call()` is back. Importable from
`nullrun.business_impact`. The constructor emits exactly what
the backend's internally-tagged serde produces, including
`extractor_id` and `extractor_version` (the backend has no
`skip_serializing_if` on these, so they are inside the hashed
bytes and must be present here too). The shared fixture digest
is byte-identical to the backend's
`DIGEST_FIXTURE_HEX_TOOL_CALL` (`9975a8b7…6ed0526966a6`).
2. The v3 `/track` single path raises typed enforcement rejections
instead of dropping them. A 422 `CONSUME_OVERBUDGET` used to be
a WARNING log line and a return value of
`TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0}` —
the call was treated as successful at the agent layer. It now
raises `NullRunConsumeOverbudgetError` carrying `reserved_cents`,
`actual_cost_cents`, `max_allowed_cents` and `epsilon_cents`.
Network errors and 5xx that name no enforcement failure still
drop and log; widening the raise to those would freeze the agent
loop on a dead backend, which is the failure mode the fail-OPEN
rows exist to prevent.
3. The /gate pre-flight now types its refusal instead of assuming
budget. A rate-limit block (NR-R002), a tool-block (NR-T003) and
a circuit-breaker trip (NR-B010) used to reach the caller as
`NullRunBudgetError` (NR-B004) — the operator reads NR-B004 as
"raise the spend cap", the actual cause is a throttle policy.
The dispatcher routes by wire code: `RATE_LIMIT_EXCEEDED` →
`NullRunRateLimitError`, `TOOL_BLOCKED` → `NullRunToolBlockedError`,
`CIRCUIT_BREAKER_TRIPPED` → typed breaker class. A response with
no machine-readable code still raises `NullRunBudgetError` (the
legacy tier is pinned). `cost_limit_exceeded` is bumped only for
`NullRunBudgetError`, so a rate-limit block no longer over-counts
the spend cap.
Carried over from 0.20.0 and still true on 0.21.0 — same class of
break, same shape of fix:
- An unclassifiable refusal now raises where 0.19.0 let the call
proceed. `NullRunUnclassifiedRefusalError` is a sibling of
`NullRunTransportError` (not a subclass), so an existing
`except NullRunTransportError:` arm will not catch it.
- `NullRunRuntime.execute(..., mode="inline")` is gone.
- `register_strict_mode_forced` / `is_strict_mode_forced` /
`@guarded` / `nullrun.handle` / `nullrun.status()` /
`nullrun.auto_instrument` are all gone (last touched in
0.18.5–0.20.0).
## Security
- **ADR-065 (DEF-TC14-002)** — `@protect` binds approvals to
nothing. Fix in five steps: (1) `BusinessImpact.tool_call()` is
restored with `extractor_id` / `extractor_version`; (2) the
envelope is carried on the call context as a contextvar, the same
home `set_call_context` uses for the model and the tool list;
(3) `Transport.check` reads the envelope and sends it alongside
the digest (the allowlist builder was dropping it before);
(4) `@protect` builds the envelope before the `/gate` pre-flight
— building it after means the two calls would carry different
envelopes, and the backend would refuse every re-entry as a side
effect; (5) a build failure (non-ASCII tool name, > 128 bytes)
degrades to `no_impact()` and logs, so the tool still runs while
the approval carries no trust binding and the server refuses the
re-entry. (`4320ac0` + `e3176d9` + `015407b` + `88ddec8` + `d11eb78`)
- **DEF-TC29-001** — `Transport.check` was dropping `tool_class`
and `mcp_annotations` from the `/gate` body. The MCP integration
had been computing them for the call context since it landed,
but the transport's explicit allowlist builder did not include
them, so a destructive MCP tool arrived at the gate as
`tool_class=None, mcp_annotations=None` — the negative case the
backend pins, not the positive case the public
`set_mcp_tool_context` API implied. Now sent unconditionally
when set; absent means "unknown", not "false". (`b64dc8a`)
- **DEF-TC4-001** — `/gate` pre-flight was raising
`NullRunBudgetError` for every refusal. The pre-flight now
routes through `_build_block_exception` and resolves the wire
code in the same order the backend resolves the HTTP status:
`details["error_code"]`, then the top-level `error_code` that
`Transport.check` already copies onto its 4xx dict, then
`explanation`. Two backend codes that had drifted out of the SDK
catalog (`BUDGET_WORKFLOW_BLOCKED`, `402`; `BUDGET_CACHE_EXCEEDED`,
`402`) are registered in the companion commit; the backend
logged `BUDGET_WORKFLOW_BLOCKED` x389 in production before that
registration, so the wire had been answering questions the SDK
could not classify. (`4ec5460` + `76efa7b`)
- **DEF-TC6-006** — `_route_track` was wrapping
`transport.track_single` in a bare `except Exception` that
logged at WARNING and returned. The transport layer had already
classified the response — a 422 `CONSUME_OVERBUDGET` becomes a
typed `NullRunConsumeOverbudgetError` — and the catch discarded
it. The drop-and-log policy the catch implements is the one the
ADR-008 table states for the `/track` batch path, a NETWORK
error; the v3 single path has no such row. The fix re-raises
`NullRunDecision` after the existing cache invalidation and
telemetry, so the blast-radius mitigation
(`DEF-CACHE-STALE-ALLOW-AFTER-OVERBUDGET`) is not traded away
for the reporting fix. (`ea7c9ee`)
## Fixed
- **DEF-TC21-001** — `WorkflowKilledInterrupt` was documented as
`BaseException`-only in three places (`docs/errors/NR-W002.md`,
`src/nullrun/breaker/exceptions.py`'s class catalog, the
`NullRunError` docstring), but the class has been an `Exception`
subclass since 0.16.6's `BreakerError` reparenting
(`9877c34`). Only the prose was wrong. `docs/errors/NR-W002.md`
also pointed at `docs/kill-contract.md` §6, a file that does
not exist. `tests/test_exception_hierarchy.py` had the same
disease: the test was named
`test_killed_interrupt_does_not_inherit_from_exception` while
asserting `issubclass(WorkflowKilledInterrupt, Exception)`.
Renamed. (`6108308`)
- **DEF-TC6-005** — `status().ws_connected` was structurally
pinned to `None`. `WebSocketConnection` has an `_running` flag
(set in `_connect`, cleared by the receive loop's `finally`);
the SDK was reading `is_open` via `getattr(..., None)`. `is_open`
appears exactly once in the SDK: on the reading side, with no
writer, no test and no producer — so the `getattr` default
fired on every call and the three states
(never-established / live / dropped) collapsed to one. The
fourth test in the new file asserts that the attribute `status()`
reads exists on a really-constructed `WebSocketConnection` AND
that `is_open` does not — a stubbed connection cannot catch it,
because the stub would carry whatever attribute the test author
assumed. (`9ccf168`)
## Tooling
- `4bbb54e` — `style: clear the three lint errors this branch's
own test files carry`. Two F841 in
`tests/test_protect_approval_roundtrip.py` (the `rt = make_runtime()`
binding was never used; `@protect` resolves the runtime from the
context, not the local) plus one I001 (DIGEST_PREFIX is upper-case
so the import sort moves it ahead of BusinessImpact). ruff CI
runs over `src/` AND `tests/`, so all three would have failed a
build.
## Verification
| Check | Result |
|---|---|
| `ruff check src tests` | All checks passed |
| `mypy src/nullrun` | Success: no issues found in 37 source files |
| `pytest -q` | **1694 passed, 1 skipped** in 109.64s (vs baseline 1597 / 1 — **+97 new tests**) |
| Scratch diff | clean (no `dist_local/`, no `*.defect*`) |
| `nullrun.__version__` | `0.21.0` |
| Wire-format | additive on `/gate` (carries `business_impact`, `tool_class`, `mcp_annotations` when set; absent means "unknown", not "false"); non-additive on `/track` (the v3 single path now raises typed enforcement rejections where it dropped them — caller-observable) |
## Commits included
```
4bbb54e style: clear the three lint errors this branch's own test files carry
d11eb78 test(protect): pin the unbuildable-envelope degradation
88ddec8 fix(protect): build the tool_call envelope before the /gate pre-flight
015407b feat(gate): send the context envelope at /gate instead of a sentinel
e3176d9 feat(sdk): carry one BusinessImpact envelope per logical action
4320ac0 feat(sdk): restore the tool_call BusinessImpact constructor
6108308 docs(kill): stop claiming the kill signal is BaseException-only
b64dc8a fix(mcp): forward tool class and annotations to /gate
ea7c9ee fix(track): propagate enforcement rejections from the v3 /track path
9ccf168 fix(sdk): read the attribute the WS connection actually has
76efa7b fix(sdk): type the /gate pre-flight refusal instead of assuming budget
4ec5460 fix(sdk): register the two backend budget codes in the SDK catalog
```
(The version-bump commit and the `uv.lock` stamp `0.20.0 → 0.21.0`
are folded into this release commit.)
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Cuts
v0.21.0— the SDK-side closure ofDEF-TC14-002(QA cycleRUN_ID
20261002T0826, 2026-10-02). Two QA cycles in a row foundthe same defect from two angles: 0.20.0's audit said "operator
approves an action and the agent still cannot run it", and this
cycle's TC-4 said "/gate blocks on a rate limit but the SDK raises
NullRunBudgetErrorso callers branch on the wrong cause". Bothwere the same root: the SDK was sending a constant sentinel instead
of the business impact envelope the backend was looking for, and
was classifying refusals by what the SDK's wrapper assumed rather
than by what the wire said.
The load-bearing fix is the five-step restoration of the
BusinessImpactenvelope (ADR-065). Since 0.18.5 every@protectcall sent a constant
{"kind":"none"}sentinel; a constant hashesto a constant, so the digest the backend stored on the approval row
at
/gatematched the digest it recomputed at/executefor everytool — while binding the approval to nothing at all. The operator's
approval card did not correspond to any particular action, and the
backend's refuse-the-reentry check had no data to refuse against.
This is a minor release (
0.21.0, not0.20.1). Three reasons:BusinessImpact.tool_call()constructor is restored on theSDK side — it was deleted in 0.18.5's deprecation sweep (aee8110)
along with the curated surface, with the prose claiming it was still
there;
/trackv3 single path now raises typed enforcement rejectionswhere it dropped them silently, so a
CONSUME_OVERBUDGETno longerreads as a successful call;
/gatepre-flight now routes through the typed dispatcher, soa
RATE_LIMIT_EXCEEDEDno longer reads as NR-B004 budgetexhausted.
Migration (read first if upgrading from 0.20.x)
Three things differ from 0.20.0. None are silent on a healthy
configuration, but each can be a working loop becoming a throwing
one for code that suppressed the failure before.
BusinessImpact.tool_call()is back. Importable fromnullrun.business_impact. The constructor emits exactly whatthe backend's internally-tagged serde produces, including
extractor_idandextractor_version(the backend has noskip_serializing_ifon these, so they are inside the hashedbytes and must be present here too). The shared fixture digest
is byte-identical to the backend's
DIGEST_FIXTURE_HEX_TOOL_CALL(9975a8b7…6ed0526966a6)./tracksingle path raises typed enforcementrejections instead of dropping them. A 422
CONSUME_OVERBUDGETused to be a WARNING log line and a return value of
TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0}—the call was treated as successful at the agent layer. It now
raises
NullRunConsumeOverbudgetErrorcarryingreserved_cents,actual_cost_cents,max_allowed_centsandepsilon_cents.Network errors and 5xx that name no enforcement failure still
drop and log; widening the raise to those would freeze the agent
loop on a dead backend, which is the failure mode the fail-OPEN
rows exist to prevent.
assuming budget. A rate-limit block (NR-R002), a tool-block
(NR-T003) and a circuit-breaker trip (NR-B010) used to reach
the caller as
NullRunBudgetError(NR-B004) — the operatorreads NR-B004 as "raise the spend cap", the actual cause is a
throttle policy. The dispatcher routes by wire code:
RATE_LIMIT_EXCEEDED→NullRunRateLimitError,TOOL_BLOCKED→NullRunToolBlockedError,CIRCUIT_BREAKER_TRIPPED→ typed breaker class. A responsewith no machine-readable code still raises
NullRunBudgetError(the legacy tier is pinned).
cost_limit_exceededis bumpedonly for
NullRunBudgetError, so a rate-limit block no longerover-counts the spend cap.
Carried over from 0.20.0 and still true on 0.21.0 — same class of
break, same shape of fix:
call proceed.
NullRunUnclassifiedRefusalErroris asibling of
NullRunTransportError(not a subclass), so anexisting
except NullRunTransportError:arm will not catchit.
NullRunRuntime.execute(..., mode="inline")is gone.register_strict_mode_forced/is_strict_mode_forced/@guarded/nullrun.handle/nullrun.status()/nullrun.auto_instrumentare all gone (last touched in0.18.5–0.20.0).
Security
@protectbinds approvals tonothing. Since 0.18.5 every
@protectcall sent a constant{"kind":"none"}sentinel; a constant hashes to a constant,so the digest the backend stored on the approval row at
/gatematched the digest it recomputed at
/executefor every tool —while binding the approval to nothing at all. The fix is in five
steps:
(1)
BusinessImpact.tool_call()is restored withextractor_id/extractor_version;(2) the envelope is carried on the call context as a contextvar,
the same home
set_call_contextuses for the model and the toollist;
(3)
Transport.checkreads the envelope and sends it alongsidethe digest (the allowlist builder was dropping it before);
(4)
@protectbuilds the envelope before the/gatepre-flight— building it after means the two calls would carry different
envelopes, and the backend would refuse every re-entry as a
side effect;
(5) a build failure (non-ASCII tool name, > 128 bytes) degrades
to
no_impact()and logs, so the tool still runs while theapproval carries no trust binding and the server refuses the
re-entry. That is a deliberate fail-OPEN on metadata — raising
out of the decorator would take down a tool call over a metadata
field. (
4320ac0+e3176d9+015407b+88ddec8+d11eb78)Transport.checkwas droppingtool_classand
mcp_annotationsfrom the/gatebody. The MCP integrationhad been computing them for the call context since it landed,
but the transport's explicit allowlist builder did not include
them, so a destructive MCP tool arrived at the gate as
tool_class=None, mcp_annotations=None— the negative case thebackend pins, not the positive case the public
set_mcp_tool_contextAPI implied.effective_tool_class()falls back to name-based classification on the negative case,
so this is a dead feature with a misleading API today — but the
day the server-side flag flips, destructive MCP tools will
silently degrade without a wire-level signal. Now sent
unconditionally when set; absent means "unknown", not "false".
(
b64dc8a)/gatepre-flight was raisingNullRunBudgetErrorfor every refusal. A rate-limit block(NR-R002), a tool-block (NR-T003) and a circuit-breaker trip
(NR-B010) all reached the caller as NR-B004 "budget exhausted",
which sends the operator looking for a spend-cap misconfiguration
when the actual cause is a throttle policy. The pre-flight now
routes through
_build_block_exceptionand resolves the wirecode in the same order the backend resolves the HTTP status:
details["error_code"], then the top-levelerror_codethatTransport.checkalready copies onto its 4xx dict, thenexplanation. The dispatcher handles all three catalog families(decision / transport / infra), not just the
NullRunBlockedExceptionone —sourceis never forwardedthrough
**detailsbecause it collides with the keyword theclass passes down itself. Two backend codes that had drifted
out of the SDK catalog (
BUDGET_WORKFLOW_BLOCKED,402;BUDGET_CACHE_EXCEEDED,402) are registered in thecompanion commit; the backend logged
BUDGET_WORKFLOW_BLOCKEDx389 in production before that registration, so the wire had
been answering questions the SDK could not classify.
(
4ec5460+76efa7b)_route_trackwas wrappingtransport.track_singlein a bareexcept Exceptionthatlogged at WARNING and returned. The transport layer had already
classified the response — a 422
CONSUME_OVERBUDGETbecomes atyped
NullRunConsumeOverbudgetError— and the catch discardedit. The drop-and-log policy the catch implements is the one
the ADR-008 table states for the
/trackbatch path, a NETWORKerror; the v3 single path has no such row. The fix re-raises
NullRunDecisionafter the existing cache invalidation andtelemetry, so the blast-radius mitigation
(
DEF-CACHE-STALE-ALLOW-AFTER-OVERBUDGET) is not traded awayfor the reporting fix. (
ea7c9ee)Fixed
WorkflowKilledInterruptwas documented asBaseException-only in three places (docs/errors/NR-W002.md,src/nullrun/breaker/exceptions.py's class catalog, theNullRunErrordocstring), but the class has been anExceptionsubclass since 0.16.6's
BreakerErrorreparenting(
9877c34). The behaviour is correct and deliberate — agentrecovery is meant to catch a kill and surface the structured
error_code/user_action;tests/test_decision_split.pydocuments the override. Only the prose was wrong, and it was
wrong in the direction that would lead the next maintainer to
revert working code.
docs/errors/NR-W002.mdalso pointed atdocs/kill-contract.md§6, a file that does not exist.tests/test_exception_hierarchy.pyhad the same disease: thetest was named
test_killed_interrupt_does_not_inherit_from_exceptionwhile asserting
issubclass(WorkflowKilledInterrupt, Exception).Renamed. (
6108308)status().ws_connectedwas structurallypinned to
None.WebSocketConnectionhas an_runningflag (set in
_connect, cleared by the receive loop'sfinally); the SDK was readingis_openviagetattr(..., None).is_openappears exactly once in the SDK: on the reading side,with no writer, no test and no producer — so the
getattrdefault fired on every call and the three states
(never-established / live / dropped) collapsed to one. The
fourth test in the new file asserts that the attribute
status()reads exists on a really-constructed
WebSocketConnectionAND that
is_opendoes not — a stubbed connection cannot catchit, because the stub would carry whatever attribute the test
author assumed. (
9ccf168)Tooling
4bbb54e—style: clear the three lint errors this branch's own test files carry. Two F841 intests/test_protect_approval_roundtrip.py(thert = make_runtime()binding was never used;
@protectresolves the runtime fromthe context, not the local) plus one I001 (DIGEST_PREFIX is
upper-case so the import sort moves it ahead of
BusinessImpact). ruff CI runs over
src/ANDtests/, so allthree would have failed a build.
Verification
ruff check src testsmypy src/nullruninstrumentation/auto.pyreproduce identically without these changes — pre-existing, out of scope)pytest -qnullrun.__version__0.21.0/gate(carriesbusiness_impact,tool_class,mcp_annotationswhen set; absent means "unknown", not "false"); non-additive on/track(the v3 single path now raises typed enforcement rejections where it dropped them — caller-observable)Commits included