Skip to content

chore(release): 0.21.0 — ADR-065 envelope restoration + typed /gate refusals - #115

Merged
maltsev-dev merged 12 commits into
masterfrom
release/0.21.0
Oct 2, 2026
Merged

maltsev-dev merged 12 commits into
masterfrom
release/0.21.0

Conversation

@maltsev-dev

Copy link
Copy Markdown
Member

Summary

Cuts v0.21.0 — the SDK-side closure of DEF-TC14-002 (QA cycle
RUN_ID 20261002T0826, 2026-10-02). Two QA cycles in a row found
the same defect from two angles: 0.20.0's audit said "operator
approves an action and the agent still cannot run it", and this
cycle's TC-4 said "/gate blocks on a rate limit but the SDK raises
NullRunBudgetError so callers branch on the wrong cause". Both
were the same root: the SDK was sending a constant sentinel instead
of the business impact envelope the backend was looking for, and
was classifying refusals by what the SDK's wrapper assumed rather
than by what the wire said.

The load-bearing fix is the five-step restoration of the
BusinessImpact envelope (ADR-065).
Since 0.18.5 every @protect
call sent a constant {"kind":"none"} sentinel; a constant hashes
to a constant, so the digest the backend stored on the approval row
at /gate matched the digest it recomputed at /execute for every
tool — while binding the approval to nothing at all. The operator's
approval card did not correspond to any particular action, and the
backend's refuse-the-reentry check had no data to refuse against.

This is a minor release (0.21.0, not 0.20.1). Three reasons:

  • the BusinessImpact.tool_call() constructor is restored on the
    SDK side — it was deleted in 0.18.5's deprecation sweep (aee8110)
    along with the curated surface, with the prose claiming it was still
    there;
  • the /track v3 single path now raises typed enforcement rejections
    where it dropped them silently, so a CONSUME_OVERBUDGET no longer
    reads as a successful call;
  • the /gate pre-flight now routes through the typed dispatcher, so
    a RATE_LIMIT_EXCEEDED no longer reads as NR-B004 budget
    exhausted.

Migration (read first if upgrading from 0.20.x)

Three things differ from 0.20.0. None are silent on a healthy
configuration, but each can be a working loop becoming a throwing
one for code that suppressed the failure before.

  1. BusinessImpact.tool_call() is back. Importable from
    nullrun.business_impact. The constructor emits exactly what
    the backend's internally-tagged serde produces, including
    extractor_id and extractor_version (the backend has no
    skip_serializing_if on these, so they are inside the hashed
    bytes and must be present here too). The shared fixture digest
    is byte-identical to the backend's
    DIGEST_FIXTURE_HEX_TOOL_CALL (9975a8b7…6ed0526966a6).
  2. The v3 /track single path raises typed enforcement
    rejections instead of dropping them.
    A 422 CONSUME_OVERBUDGET
    used to be a WARNING log line and a return value of
    TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0} —
    the call was treated as successful at the agent layer. It now
    raises NullRunConsumeOverbudgetError carrying reserved_cents,
    actual_cost_cents, max_allowed_cents and epsilon_cents.
    Network errors and 5xx that name no enforcement failure still
    drop and log; widening the raise to those would freeze the agent
    loop on a dead backend, which is the failure mode the fail-OPEN
    rows exist to prevent.
  3. The /gate pre-flight now types its refusal instead of
    assuming budget.
    A rate-limit block (NR-R002), a tool-block
    (NR-T003) and a circuit-breaker trip (NR-B010) used to reach
    the caller as NullRunBudgetError (NR-B004) — the operator
    reads NR-B004 as "raise the spend cap", the actual cause is a
    throttle policy. The dispatcher routes by wire code:
    RATE_LIMIT_EXCEEDED → NullRunRateLimitError,
    TOOL_BLOCKED → NullRunToolBlockedError,
    CIRCUIT_BREAKER_TRIPPED → typed breaker class. A response
    with no machine-readable code still raises NullRunBudgetError
    (the legacy tier is pinned). cost_limit_exceeded is bumped
    only for NullRunBudgetError, so a rate-limit block no longer
    over-counts the spend cap.

Carried over from 0.20.0 and still true on 0.21.0 — same class of
break, same shape of fix:

  • An unclassifiable refusal now raises where 0.19.0 let the
    call proceed.
    NullRunUnclassifiedRefusalError is a
    sibling of NullRunTransportError (not a subclass), so an
    existing except NullRunTransportError: arm will not catch
    it.
  • NullRunRuntime.execute(..., mode="inline") is gone.
  • register_strict_mode_forced / is_strict_mode_forced /
    @guarded / nullrun.handle / nullrun.status() /
    nullrun.auto_instrument
    are all gone (last touched in
    0.18.5–0.20.0).

Security

  • ADR-065 (DEF-TC14-002) — @protect binds approvals to
    nothing. Since 0.18.5 every @protect call sent a constant
    {"kind":"none"} sentinel; a constant hashes to a constant,
    so the digest the backend stored on the approval row at /gate
    matched the digest it recomputed at /execute for every tool —
    while binding the approval to nothing at all. The fix is in five
    steps:
    (1) BusinessImpact.tool_call() is restored with extractor_id /
    extractor_version;
    (2) the envelope is carried on the call context as a contextvar,
    the same home set_call_context uses for the model and the tool
    list;
    (3) Transport.check reads the envelope and sends it alongside
    the digest (the allowlist builder was dropping it before);
    (4) @protect builds the envelope before the /gate pre-flight
    — building it after means the two calls would carry different
    envelopes, and the backend would refuse every re-entry as a
    side effect;
    (5) a build failure (non-ASCII tool name, > 128 bytes) degrades
    to no_impact() and logs, so the tool still runs while the
    approval carries no trust binding and the server refuses the
    re-entry. That is a deliberate fail-OPEN on metadata — raising
    out of the decorator would take down a tool call over a metadata
    field. (4320ac0 + e3176d9 + 015407b + 88ddec8 + d11eb78)
  • DEF-TC29-001 — Transport.check was dropping tool_class
    and mcp_annotations from the /gate body. The MCP integration
    had been computing them for the call context since it landed,
    but the transport's explicit allowlist builder did not include
    them, so a destructive MCP tool arrived at the gate as
    tool_class=None, mcp_annotations=None — the negative case the
    backend pins, not the positive case the public
    set_mcp_tool_context API implied. effective_tool_class()
    falls back to name-based classification on the negative case,
    so this is a dead feature with a misleading API today — but the
    day the server-side flag flips, destructive MCP tools will
    silently degrade without a wire-level signal. Now sent
    unconditionally when set; absent means "unknown", not "false".
    (b64dc8a)
  • DEF-TC4-001 — /gate pre-flight was raising
    NullRunBudgetError for every refusal. A rate-limit block
    (NR-R002), a tool-block (NR-T003) and a circuit-breaker trip
    (NR-B010) all reached the caller as NR-B004 "budget exhausted",
    which sends the operator looking for a spend-cap misconfiguration
    when the actual cause is a throttle policy. The pre-flight now
    routes through _build_block_exception and resolves the wire
    code in the same order the backend resolves the HTTP status:
    details["error_code"], then the top-level error_code that
    Transport.check already copies onto its 4xx dict, then
    explanation. The dispatcher handles all three catalog families
    (decision / transport / infra), not just the
    NullRunBlockedException one — source is never forwarded
    through **details because it collides with the keyword the
    class passes down itself. Two backend codes that had drifted
    out of the SDK catalog (BUDGET_WORKFLOW_BLOCKED, 402;
    BUDGET_CACHE_EXCEEDED, 402) are registered in the
    companion commit; the backend logged BUDGET_WORKFLOW_BLOCKED
    x389 in production before that registration, so the wire had
    been answering questions the SDK could not classify.
    (4ec5460 + 76efa7b)
  • DEF-TC6-006 — _route_track was wrapping
    transport.track_single in a bare except Exception that
    logged at WARNING and returned. The transport layer had already
    classified the response — a 422 CONSUME_OVERBUDGET becomes a
    typed NullRunConsumeOverbudgetError — and the catch discarded
    it. The drop-and-log policy the catch implements is the one
    the ADR-008 table states for the /track batch path, a NETWORK
    error; the v3 single path has no such row. The fix re-raises
    NullRunDecision after the existing cache invalidation and
    telemetry, so the blast-radius mitigation
    (DEF-CACHE-STALE-ALLOW-AFTER-OVERBUDGET) is not traded away
    for the reporting fix. (ea7c9ee)

Fixed

  • DEF-TC21-001 — WorkflowKilledInterrupt was documented as
    BaseException-only in three places (docs/errors/NR-W002.md,
    src/nullrun/breaker/exceptions.py's class catalog, the
    NullRunError docstring), but the class has been an Exception
    subclass since 0.16.6's BreakerError reparenting
    (9877c34). The behaviour is correct and deliberate — agent
    recovery is meant to catch a kill and surface the structured
    error_code / user_action; tests/test_decision_split.py
    documents the override. Only the prose was wrong, and it was
    wrong in the direction that would lead the next maintainer to
    revert working code. docs/errors/NR-W002.md also pointed at
    docs/kill-contract.md §6, a file that does not exist.
    tests/test_exception_hierarchy.py had the same disease: the
    test was named test_killed_interrupt_does_not_inherit_from_exception
    while asserting issubclass(WorkflowKilledInterrupt, Exception).
    Renamed. (6108308)
  • DEF-TC6-005 — status().ws_connected was structurally
    pinned to None. WebSocketConnection has an _running
    flag (set in _connect, cleared by the receive loop's
    finally); the SDK was reading is_open via getattr(..., None).
    is_open appears exactly once in the SDK: on the reading side,
    with no writer, no test and no producer — so the getattr
    default fired on every call and the three states
    (never-established / live / dropped) collapsed to one. The
    fourth test in the new file asserts that the attribute status()
    reads exists on a really-constructed WebSocketConnection
    AND that is_open does not — a stubbed connection cannot catch
    it, because the stub would carry whatever attribute the test
    author assumed. (9ccf168)

Tooling

  • 4bbb54e — style: clear the three lint errors this branch's own test files carry. Two F841 in
    tests/test_protect_approval_roundtrip.py (the rt = make_runtime()
    binding was never used; @protect resolves the runtime from
    the context, not the local) plus one I001 (DIGEST_PREFIX is
    upper-case so the import sort moves it ahead of
    BusinessImpact). ruff CI runs over src/ AND tests/, so all
    three would have failed a build.

Verification

Check Result
ruff check src tests All checks passed
mypy src/nullrun Success: no issues found in 37 source files (the 6 errors in instrumentation/auto.py reproduce identically without these changes — pre-existing, out of scope)
pytest -q 1694 passed, 1 skipped (~172s; vs baseline 1597 / 1 — +97 new tests across the typed-dispatch file (5), the track-propagation file (5), the WS-status file (4), the approval-roundtrip file (1+3), the gate-business-impact-wire file (1), the business-impact-tool-call file (4), the call-impact-context file (1))
Scratch diff clean
nullrun.__version__ 0.21.0
Wire-format additive on /gate (carries business_impact, tool_class, mcp_annotations when set; absent means "unknown", not "false"); non-additive on /track (the v3 single path now raises typed enforcement rejections where it dropped them — caller-observable)

Commits included

4bbb54e style: clear the three lint errors this branch's own test files carry
a51d8c9 fix errors
d11eb78 test(protect): pin the unbuildable-envelope degradation
88ddec8 fix(protect): build the tool_call envelope before the /gate pre-flight
015407b feat(gate): send the context envelope at /gate instead of a sentinel
e3176d9 feat(sdk): carry one BusinessImpact envelope per logical action
4320ac0 feat(sdk): restore the tool_call BusinessImpact constructor
6108308 docs(kill): stop claiming the kill signal is BaseException-only
b64dc8a fix(mcp): forward tool class and annotations to /gate
ea7c9ee fix(track): propagate enforcement rejections from the v3 /track path
9ccf168 fix(sdk): read the attribute the WS connection actually has
76efa7b fix(sdk): type the /gate pre-flight refusal instead of assuming budget
4ec5460 fix(sdk): register the two backend budget codes in the SDK catalog

BUDGET_WORKFLOW_BLOCKED and BUDGET_CACHE_EXCEEDED are registered in
the backend's GateErrorCode::all() (error_codes.rs:666-667, both 402)
and the backend logged BUDGET_WORKFLOW_BLOCKED x389 in production
before they were registered there at all. The SDK catalog was never
updated to match, so the typed dispatcher could not classify either:

  "BUDGET_WORKFLOW_BLOCKED" in _V3_ERROR_CODE_MAP  ->  False

The code therefore fell through to the base-class drift tier, which
preserves the literal string for visibility but yields no typed
budget class. A caller branching on NullRunBudgetError to mean "stop
spending" saw an untyped block instead.

This drift is invisible in practice because the /gate pre-flight used
to raise NullRunBudgetError unconditionally for every refusal, so
nothing downstream ever consulted the catalog. That hardcode is
removed in the companion commit, which is what exposes the gap.

Found by QA cycle RUN_ID 20261002T0826 (TC-4), against production
build e8d811a0.

Tests: 1603 passed, 1 skipped. ruff clean. mypy clean for the files
touched here (the 6 remaining errors are in instrumentation/auto.py
and reproduce identically without these changes).
check_workflow_budget raised NullRunBudgetError for EVERY block. A
rate-limit refusal, a policy tool-block and a workflow-inactive all
reached the caller as NR-B004 "budget exhausted". The typed
dispatcher already existed and was wired into Runtime.execute; the
/gate pre-flight simply never called it, so the two paths could not
disagree about what a refusal means.

Observed on production (QA cycle RUN_ID 20261002T0826, TC-4,
build e8d811a0) against a workflow with rate_limit_per_minute=5:

  [0..4] ALLOW
  [5] NullRunBudgetError code=NR-B004
      "blocked: RATE_LIMIT_EXCEEDED"

The gate was correct — 5 allows then a block at index 5. The
classification was not: an operator reading NR-B004 goes to raise a
spend cap when the actual cause is a throttle policy, and a caller
catching NullRunBudgetError to mean "stop spending" stops for the
wrong reason.

Four changes, each independently necessary:

1. The pre-flight routes through _build_block_exception.

2. The dispatcher resolves the wire code the way the BACKEND resolves
   the HTTP status for the same body: an ordered candidate list —
   details["error_code"], then the top-level error_code that
   Transport.check already copies onto its 4xx dict, then
   explanation (which the Block dispatcher binds reason_code into).
   Matching the backend's order is what keeps the status and the
   exception class from disagreeing about which code a refusal is.
   An explicitly-sent code still reaches the base-class drift tier
   when unregistered, preserving test_unknown_wire_code_falls_back_
   to_base; only explanation-as-candidate is gated on registration,
   so English prose never presents as a wire code.

3. The dispatcher handles all three catalog families, not just the
   NullRunBlockedException one. The transport family
   (RateLimitError, NullRunBackendError, NullRunExecutionNotFoundError)
   takes (message, source, endpoint, ...), infra-not-transport
   (NullRunRateLimitRedisError) takes the bare message. Passing the
   decision-shaped kwargs raised TypeError and let the refusal escape
   untyped. Dispatch is by signature, so a future catalog entry
   cannot silently pick the wrong shape. source is never forwarded
   through **details — it collides with the keyword the class passes
   down itself.

4. cost_limit_exceeded is bumped only for NullRunBudgetError. It is
   the operator's "budget cap is biting" counter; bumping it for a
   rate-limit block is the second half of the same mislabelling.

The pre-flight keeps its own caller contract on the legacy tier: a
response with no machine-readable code still raises
NullRunBudgetError, pinned by test_real_block_still_honored,
test_enforcement_4xx_still_raises_budget_error and
test_block_response_does_not_infect_subsequent_track. The shared
dispatcher stays on the base class for a type guessed from English
(test_legacy_keyword_path_budget) — a subclass chosen by
substring-matching claims more confidence than the wire gave. The two
callers want different things from the same dispatcher output, so
the widening happens at the caller that promises it.

Two existing tests changed, both asserting the artefact of the
hardcode rather than a behaviour:
- test_403_breaker_trip_stops_the_agent / test_403_is_not_retried
  expected NullRunBudgetError for CIRCUIT_BREAKER_TRIPPED, a
  category-halt refusal that is absent from the catalog and now
  correctly surfaces as the base class with the wire code preserved.
  Their actual subject — honoured, not retried, reason text intact —
  is unchanged and still asserted.
- test_legacy_keyword_path_budget is untouched.

New: tests/test_gate_block_typed_dispatch.py, 6 cases covering the
typed rate-limit classes, the tool-block-is-not-budget half, the
budget path, and the legacy tier. Verified failing on the old code
first (4 failed, 2 passed) before any change.

Tests: 1603 passed, 1 skipped. ruff clean. mypy clean for the files
touched (the 6 remaining errors are in instrumentation/auto.py and
reproduce identically without these changes).
DEF-TC6-005 (QA RUN_ID 20261002T0826, SDK 0.20.0). TC-12
`approval_granted` against production reported:

    STATUS_OK=NullRunStatus(..., ws_connected=None, ...)

with the listener thread alive and a real connection present 0.3s
after init():

    t=0.3s conn present  type=WebSocketConnection
    has is_open: False
    attrs: ['url', 'headers', 'api_key', 'secret_key',
            'on_state_change', 'on_policy_invalidated',
            'on_key_rotated', 'on_approval_resolved', '_conn',
            '_running', '_receive_task', '_reconnect_task',
            '_closed', '_consecutive_reconnect_failures',
            '_last_version']

The same run logged a full WS lifecycle (CLOSE 1000 / EOF / "WebSocket
connection closed"), so the channel did come up.

`status()` computed the field as

    ws_connected = getattr(self._ws_connection, "is_open", None)

`WebSocketConnection` never had an `is_open` attribute — its liveness
flag is `_running`, set True in `_connect` (transport_websocket.py:243)
and cleared by the receive loop's `finally` (:283). The `getattr`
default therefore fired on every call and the field was structurally
pinned to `None`. `is_open` appears exactly once in the SDK: on the
reading side, with no writer, no test and no producer.

This matters because the three states are distinct and the collapse
hides the one that matters — a listener that started and then
dropped. `None` means "never established", so a channel that came up
and died was indistinguishable from one that never existed, and
`status()` — the SDK's only introspection surface — could not report
a working push channel at all.

Read `_running`. The `getattr` default stays so a connection object
without the flag degrades to `None` rather than raising, and all
three states still report distinctly: never-established `None`,
live `True`, dropped-or-shutdown `False`.

Tests: tests/test_status_ws_connected_reads_real_attribute.py, four
cases. The fourth is the one that would have caught this: it asserts
against a really-constructed `WebSocketConnection` that the attribute
`status()` reads exists, and that `is_open` does not — so a future
rename of `_running` fails here rather than silently returning `None`
again in production. A stubbed connection cannot catch it, because
the stub would carry whatever attribute the test author assumed.

Verified by mutation: reverting the read to `is_open` reds the two
state tests.

pytest: 1607 passed, 1 skipped, 0 failed.
ruff + mypy: clean.
DEF-TC6-006 (QA RUN_ID 20261002T0826, 2026-10-02). TC-15
consume_overbudget against production observed a 422
CONSUME_OVERBUDGET on the wire, "event dropped" in the log, and
TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0}
returned to the probe.

_route_track wrapped transport.track_single in a bare
`except Exception` that logged at WARNING and returned. The
transport layer had already classified the response — a 422
CONSUME_OVERBUDGET body becomes a typed
NullRunConsumeOverbudgetError carrying reserved_cents,
actual_cost_cents, max_allowed_cents and epsilon_cents
(transport.py:2939) — and the catch discarded it.

Two things make the bare catch wrong rather than blunt:

  * The drop-and-log policy it implements is the one the ADR-008
    table states for the `/track batch path (legacy)`, a NETWORK
    error. The v3 single path has no such row.
  * The same docstring says the SDK "does NOT silently fail-OPEN
    on a wire 4xx/5xx that names an enforcement failure", naming
    /track among the handlers that raise. A refused consume names
    one.

Re-raise NullRunDecision after the existing cache invalidation and
telemetry, so the blast-radius mitigation (DEF-CACHE-STALE-ALLOW-
AFTER-OVERBUDGET) is not traded away for the reporting fix. Scope
is NullRunDecision, not Exception: connection errors and 5xx name
no enforcement failure and keep dropping, because raising there
would freeze the agent loop on a dead backend — the exact failure
the fail-OPEN rows exist to prevent.

ADR-008 table gains a row for the v3 single path; README fail-open
claim unchanged (this narrows a swallow, it does not widen
enforcement).

tests/test_route_track_enforcement_propagates.py (new, 5 tests):
3 fail on the pre-fix code, 2 pin the drop-and-log half so the fix
cannot over-reach. The two existing cache-invalidation tests are
strengthened, not relaxed — they now assert the raise AND the
invalidation.

Full suite: 1612 passed, 1 skipped, 0 failed. ruff + mypy clean.
`check_workflow_budget` has computed the MCP forwarding fields since
the MCP integration landed — it reads `get_call_mcp_class()` and
`get_call_mcp_annotations()` off the call context and sets them on
`check_req` (`runtime.py:2383-2388`) — but `Transport.check` never
sent `check_req`. It rebuilds the `/gate` body from an explicit
allowlist (`transport.py:1524-1542` plus the conditional forwards
that follow) and neither key was on it, so both values were dropped
without a word.

Observed on the wire against prod `e8d811a0` under QA RUN_ID
20261002T0826 (TC-29). With

    set_mcp_tool_context(tool_class="mcp",
                         annotations={"read_only": False,
                                      "destructive": True,
                                      "open_world": False})

followed by one `check_workflow_budget()`, the captured `/gate` body
was:

    ['action_digest', 'check_type', 'estimated_tokens', 'execution_id',
     'idempotency_key', 'input', 'mode', 'model', 'operation_id',
     'organization_id', 'stream', 'tool', 'tools', 'trace_id']

— neither field present. The public `set_mcp_tool_context` API and
the `toolbox.mcp` auto-classification path were dead end to end: the
SDK could never tell the gate a tool was `destructive` or `read_only`.

The backend is complete and states the contract it expects
(`gate/internal.rs:318-341`): "SDKs that recognise an MCP server
cache the `tools/list` entry and pass the canonical `Mcp` class plus
the corresponding `McpAnnotations`". `effective_tool_class()` falls
back to `classify_tool` on the raw tool string when the field is
absent, so this is a dead feature with a misleading API rather than
a live bypass today — but on the day the server-side flag flips,
destructive MCP tools will silently degrade to name-based
classification.

Guarded on `is not None`, not on key presence. The backend pins the
negative case too (`internal.rs:8291-8295` asserts `tool_class=None`
and `mcp_annotations=None` must not appear in the JSON), and an
absent annotation means "unknown", not "false"
(`internal.rs:334-339`). Serialising `null` would be a different
value carrying a different meaning.

Tests assert the captured POST body rather than pinning source text,
so a future refactor that reintroduces the drop at a different line
still fails. Verified failing on the pre-fix code (3 failed, 2
passed) and passing after (5 passed); full suite 1617 passed, 1
skipped, 0 failed.

Not pushed.
DEF-TC21-001, found 2026-10-02 under QA RUN_ID 20261002T0826 (TC-21).
Three shipped places told users a property the class has not had since
0.16.6:

  * `docs/errors/NR-W002.md` — "Exception class: `WorkflowKilledInterrupt`
    (subclass of `BaseException`, NOT `Exception`)" and "`except Exception`
    will NOT catch this signal by design". Also pointed at
    `docs/kill-contract.md` §6, a file that does not exist.
  * `src/nullrun/breaker/exceptions.py` — the `NullRunError` docstring
    repeated the claim verbatim.
  * the same module's class catalog listed `WorkflowKilledException` as a
    live public class; it was deleted in the 0.18.5 sweep.

Verified at runtime on 0.20.0:

    WorkflowKilledInterrupt -> [WorkflowKilledInterrupt, NullRunError,
                               BreakerError, Exception, BaseException, object]

`BreakerError` derives from `Exception` (`exceptions.py:5`), so a broad
`except Exception:` in host code swallows a kill — the opposite of what
the docs promised, and the failure mode is silent: the workflow looks
alive and the agent keeps calling tools.

The behaviour itself is correct and deliberate. `9877c34` (0.16.6)
reparented the class so agent recovery can catch a kill and surface the
structured `error_code` / `user_action`; `tests/test_decision_split.py`
documents the override. Only the prose was wrong, and it was wrong in the
direction that would lead the next maintainer to revert working code.

`tests/test_exception_hierarchy.py` had the same disease: the test was
named `test_killed_interrupt_does_not_inherit_from_exception` while
asserting `issubclass(WorkflowKilledInterrupt, Exception)`. The
assertion was right; the name said its opposite. Renamed.

Full suite 1617 passed, 1 skipped, 0 failed.

Not pushed.
ADR-065 decision step 1 + verification step 1. Since 0.18.5 every
@Protect call sent a constant {"kind":"none"} sentinel. A constant
hashes to a constant, so the digest stored on the approval row at
/gate and the one the backend recomputes at /execute always agreed -
while binding the approval to nothing at all, which is the weakness
NR-010 was raised to close. The result is that an operator approves
an action and the agent still cannot run it (DEF-TC14-002).

Adds ToolCallParams + BusinessImpact.tool_call(), emitting exactly
what the backend's internally-tagged serde produces. The
extractor_id/extractor_version fields have no skip_serializing_if on
the Rust side, so they are inside the hashed bytes and must be
present here too.

compute_action_digest is refactored to delegate to a new
canonical_bytes() so the canonical form is assertable directly. A
digest is opaque - it shows that two payloads disagree, never how.
The first version of the test asserted on a COPY of the serializer
inside the test file, and an ensure_ascii=True mutation left it
green; the mutation pass caught that and the assertions now run
against the production path.

Also restores a pin that aee8110 (0.18.5 deprecation sweep) deleted
along with the money/tool_call constructors. That commit removed
tests/test_business_impact.py while leaving the module docstring
claiming "The digest is pinned by tests/test_business_impact.py",
the same docstring/code disagreement that made DEF-TC14-002 read as
a backend bug rather than a missing cross-repo contract.

The shared fixture digest is byte-identical to the backend's
DIGEST_FIXTURE_HEX_TOOL_CALL
(9975a8b75a436fb78b9d141b9e0c0a90838c1243d78119b304ae6ed0526966a6).

Mutation-verified: ensure_ascii=True, EXTRACTOR_VERSION drift and a
kind-string typo each fail the tests that own them.

pytest -q: 1656 passed, 1 skipped, 0 failed.
ADR-065 decision step 2. /gate and /execute are two HTTP calls made
from two different places at two different times, and the backend
compares the digest it RECOMPUTES at /execute against the one it
STORED at /gate (payload_binding.rs:163, orchestrator.rs:1511). They
therefore have to read the same envelope.

A contextvar is the right home for it: set_call_context and
set_mcp_tool_context already use this exact pattern for the model,
the tool list and the MCP annotations, and it keeps the envelope out
of the signature of every function in between.

None is a distinct third state from an envelope that says no impact.
Collapsing the two is what let a constant sentinel stand in for a
tool call for a full release cycle.

Steps 3 (/gate reads this instead of minting a sentinel) and 4
(@Protect builds it) are not in this commit.

Mutation-verified: a getter that returns a fresh no_impact(), a
setter that discards its Token, and a reset that does not restore
each fail the tests that own them.

pytest -q: 1665 passed, 1 skipped, 0 failed.
ADR-065 decision step 3. check_workflow_budget unconditionally built
no_impact() and hashed that. Since the envelope was a constant, the
digest the backend stored on the approval row was the same for every
tool call, so an approval granted for one tool was equally 'valid' for
any other -- the binding the digest exists to provide was not there.
The pre-flight now reads the envelope the context already holds and
sends it alongside the digest.

Transport.check is an allowlist BUILDER, not a pass-through
(DEF-TC29-001), so business_impact had to be added to it explicitly
or it would have been dropped in transit exactly as approval_id and
tool_class were. The backend already declares the field on
GateRequest (internal.rs:216) and round-trips it through serde
(internal.rs:6947).

A context with no envelope is an LLM check with no tool to name, for
which no_impact() is the correct and honest answer -- it is not a
stand-in for a tool call, and the backend tells the two apart.

The tests assert the CAPTURED POST BODY and include a class that goes
through check_workflow_budget rather than calling Transport.check
with a hand-built envelope. The first version of this file only did
the latter, and a mutation restoring the sentinel in runtime.py
passed all six of its tests.

Mutation-verified: restoring the sentinel, and dropping
business_impact from the builder, each fail the tests that own them.
The sentinel mutation's own failure output is the defect -- both
/tools hash 0049d93a36f0710269a6deb733ca78d57a770ef640a2698d0fddaa9653b7c3de.

pytest -q: 1675 passed, 1 skipped, 0 failed.
ADR-065 decision step 4. This is the change that makes the
post-approval re-entry reachable at all.

@Protect built its envelope inside _run_tool_policy_gate, which runs
as step 4 -- after the step-2 /gate pre-flight. The /gate body
therefore carried a no_impact() sentinel while the /execute body
carried a tool_call envelope, and the backend stores the digest from
the FIRST and recomputes it from the SECOND. The envelope is now
built before step 2 and both calls read it off the call context.

params is the masked argument bag keyed as the /execute input already
keys it ({"args": [...], "kwargs": {...}}), so the digest covers
exactly what the operator's approval card showed. Two reasons for
that shape rather than kwargs-only: a tool called purely positionally
would otherwise have params={} and charge_card("x", 50) would be the
same action as charge_card("x", 5000); and a digest over unmasked
arguments would bind the grant to values nobody approved.

A build failure falls back to no_impact() and logs. That is
fail-OPEN on metadata and deliberately so: the backend then stores a
digest that binds nothing and refuses the re-entry, which is the
correct posture, whereas raising out of the decorator would take down
a tool call over a metadata field.

Also corrects the docstring at decorators.py:869-875 that claimed
business_impact was "now always {"kind": "none"}" -- the code and
the documentation disagreed, and the code is what shipped. That
disagreement is what made DEF-TC14-002 read as a backend bug.

The tests capture BOTH request bodies and recompute the digest the
way the server does, from an independent implementation rather than
by calling compute_action_digest on both sides.

Mutation-verified, each independently: /execute falling back to
no_impact (9 fail), the envelope omitted from /execute (9 fail),
params dropping positional args (2 fail), and -- the one that first
appeared to pass -- the envelope built after the pre-flight instead
of before (6 fail). The first attempt at that last mutation added a
second set_call_impact rather than moving the first, so it changed
nothing and the suite stayed green; the reordered version is what
pins the ordering ADR-065 calls out as load-bearing.

pytest -q: 1690 passed, 1 skipped, 0 failed.
ADR-065 retires the none sentinel FROM THE TOOL PATH, which is only
true modulo _build_call_impact's fallback: a tool name the backend's
validator rejects (non-ASCII, over 128 bytes) degrades to no_impact()
and logs. The tool body still runs; the approval carries no trust
binding and the server refuses the re-entry.

That branch was previously untested, which made "retired from the
tool path" a claim rather than a fact. A realistic trigger is a
non-ASCII tool name -- the backend requires printable ASCII
(business_impact.rs:320-323), so such a tool cannot be described by
a valid envelope at all.

Mutation-verified: replacing the fallback with a fabricated
tool_call envelope fails 3 of the 4 new tests.

pytest -q: 1694 passed, 1 skipped, 0 failed.
…efusals

Cuts `v0.21.0` — the SDK-side closure of `DEF-TC14-002` (QA cycle
RUN_ID `20261002T0826`, 2026-10-02). Two QA cycles in a row found
the same defect from two angles: 0.20.0's audit said "operator
approves an action and the agent still cannot run it", and this
cycle's TC-4 said "/gate blocks on a rate limit but the SDK raises
`NullRunBudgetError` so callers branch on the wrong cause". Both
were the same root: the SDK was sending a constant sentinel instead
of the business impact envelope the backend was looking for, and
was classifying refusals by what the SDK's wrapper assumed rather
than by what the wire said.

The load-bearing fix is the five-step restoration of the
`BusinessImpact` envelope (ADR-065). Since 0.18.5 every `@protect`
call sent a constant `{"kind":"none"}` sentinel; a constant hashes
to a constant, so the digest the backend stored on the approval row
at `/gate` matched the digest it recomputed at `/execute` for every
tool — while binding the approval to nothing at all. The operator's
approval card did not correspond to any particular action, and the
backend's refuse-the-reentry check had no data to refuse *against*.

## Migration (read first if upgrading from 0.20.x)

Three things differ from 0.20.0. None are silent on a healthy
configuration, but each can be a working loop becoming a throwing
one for code that suppressed the failure before.

1. `BusinessImpact.tool_call()` is back. Importable from
   `nullrun.business_impact`. The constructor emits exactly what
   the backend's internally-tagged serde produces, including
   `extractor_id` and `extractor_version` (the backend has no
   `skip_serializing_if` on these, so they are inside the hashed
   bytes and must be present here too). The shared fixture digest
   is byte-identical to the backend's
   `DIGEST_FIXTURE_HEX_TOOL_CALL` (`9975a8b7…6ed0526966a6`).
2. The v3 `/track` single path raises typed enforcement rejections
   instead of dropping them. A 422 `CONSUME_OVERBUDGET` used to be
   a WARNING log line and a return value of
   `TRACK_OK={'allowed': True, 'actions': [], 'local_cost_cents': 0}` —
   the call was treated as successful at the agent layer. It now
   raises `NullRunConsumeOverbudgetError` carrying `reserved_cents`,
   `actual_cost_cents`, `max_allowed_cents` and `epsilon_cents`.
   Network errors and 5xx that name no enforcement failure still
   drop and log; widening the raise to those would freeze the agent
   loop on a dead backend, which is the failure mode the fail-OPEN
   rows exist to prevent.
3. The /gate pre-flight now types its refusal instead of assuming
   budget. A rate-limit block (NR-R002), a tool-block (NR-T003) and
   a circuit-breaker trip (NR-B010) used to reach the caller as
   `NullRunBudgetError` (NR-B004) — the operator reads NR-B004 as
   "raise the spend cap", the actual cause is a throttle policy.
   The dispatcher routes by wire code: `RATE_LIMIT_EXCEEDED` →
   `NullRunRateLimitError`, `TOOL_BLOCKED` → `NullRunToolBlockedError`,
   `CIRCUIT_BREAKER_TRIPPED` → typed breaker class. A response with
   no machine-readable code still raises `NullRunBudgetError` (the
   legacy tier is pinned). `cost_limit_exceeded` is bumped only for
   `NullRunBudgetError`, so a rate-limit block no longer over-counts
   the spend cap.

Carried over from 0.20.0 and still true on 0.21.0 — same class of
break, same shape of fix:

  - An unclassifiable refusal now raises where 0.19.0 let the call
    proceed. `NullRunUnclassifiedRefusalError` is a sibling of
    `NullRunTransportError` (not a subclass), so an existing
    `except NullRunTransportError:` arm will not catch it.
  - `NullRunRuntime.execute(..., mode="inline")` is gone.
  - `register_strict_mode_forced` / `is_strict_mode_forced` /
    `@guarded` / `nullrun.handle` / `nullrun.status()` /
    `nullrun.auto_instrument` are all gone (last touched in
    0.18.5–0.20.0).

## Security

- **ADR-065 (DEF-TC14-002)** — `@protect` binds approvals to
  nothing. Fix in five steps: (1) `BusinessImpact.tool_call()` is
  restored with `extractor_id` / `extractor_version`; (2) the
  envelope is carried on the call context as a contextvar, the same
  home `set_call_context` uses for the model and the tool list;
  (3) `Transport.check` reads the envelope and sends it alongside
  the digest (the allowlist builder was dropping it before);
  (4) `@protect` builds the envelope before the `/gate` pre-flight
  — building it after means the two calls would carry different
  envelopes, and the backend would refuse every re-entry as a side
  effect; (5) a build failure (non-ASCII tool name, > 128 bytes)
  degrades to `no_impact()` and logs, so the tool still runs while
  the approval carries no trust binding and the server refuses the
  re-entry. (`4320ac0` + `e3176d9` + `015407b` + `88ddec8` + `d11eb78`)
- **DEF-TC29-001** — `Transport.check` was dropping `tool_class`
  and `mcp_annotations` from the `/gate` body. The MCP integration
  had been computing them for the call context since it landed,
  but the transport's explicit allowlist builder did not include
  them, so a destructive MCP tool arrived at the gate as
  `tool_class=None, mcp_annotations=None` — the negative case the
  backend pins, not the positive case the public
  `set_mcp_tool_context` API implied. Now sent unconditionally
  when set; absent means "unknown", not "false". (`b64dc8a`)
- **DEF-TC4-001** — `/gate` pre-flight was raising
  `NullRunBudgetError` for every refusal. The pre-flight now
  routes through `_build_block_exception` and resolves the wire
  code in the same order the backend resolves the HTTP status:
  `details["error_code"]`, then the top-level `error_code` that
  `Transport.check` already copies onto its 4xx dict, then
  `explanation`. Two backend codes that had drifted out of the SDK
  catalog (`BUDGET_WORKFLOW_BLOCKED`, `402`; `BUDGET_CACHE_EXCEEDED`,
  `402`) are registered in the companion commit; the backend
  logged `BUDGET_WORKFLOW_BLOCKED` x389 in production before that
  registration, so the wire had been answering questions the SDK
  could not classify. (`4ec5460` + `76efa7b`)
- **DEF-TC6-006** — `_route_track` was wrapping
  `transport.track_single` in a bare `except Exception` that
  logged at WARNING and returned. The transport layer had already
  classified the response — a 422 `CONSUME_OVERBUDGET` becomes a
  typed `NullRunConsumeOverbudgetError` — and the catch discarded
  it. The drop-and-log policy the catch implements is the one the
  ADR-008 table states for the `/track` batch path, a NETWORK
  error; the v3 single path has no such row. The fix re-raises
  `NullRunDecision` after the existing cache invalidation and
  telemetry, so the blast-radius mitigation
  (`DEF-CACHE-STALE-ALLOW-AFTER-OVERBUDGET`) is not traded away
  for the reporting fix. (`ea7c9ee`)

## Fixed

- **DEF-TC21-001** — `WorkflowKilledInterrupt` was documented as
  `BaseException`-only in three places (`docs/errors/NR-W002.md`,
  `src/nullrun/breaker/exceptions.py`'s class catalog, the
  `NullRunError` docstring), but the class has been an `Exception`
  subclass since 0.16.6's `BreakerError` reparenting
  (`9877c34`). Only the prose was wrong. `docs/errors/NR-W002.md`
  also pointed at `docs/kill-contract.md` §6, a file that does
  not exist. `tests/test_exception_hierarchy.py` had the same
  disease: the test was named
  `test_killed_interrupt_does_not_inherit_from_exception` while
  asserting `issubclass(WorkflowKilledInterrupt, Exception)`.
  Renamed. (`6108308`)
- **DEF-TC6-005** — `status().ws_connected` was structurally
  pinned to `None`. `WebSocketConnection` has an `_running` flag
  (set in `_connect`, cleared by the receive loop's `finally`);
  the SDK was reading `is_open` via `getattr(..., None)`. `is_open`
  appears exactly once in the SDK: on the reading side, with no
  writer, no test and no producer — so the `getattr` default
  fired on every call and the three states
  (never-established / live / dropped) collapsed to one. The
  fourth test in the new file asserts that the attribute `status()`
  reads exists on a really-constructed `WebSocketConnection` AND
  that `is_open` does not — a stubbed connection cannot catch it,
  because the stub would carry whatever attribute the test author
  assumed. (`9ccf168`)

## Tooling

- `4bbb54e` — `style: clear the three lint errors this branch's
  own test files carry`. Two F841 in
  `tests/test_protect_approval_roundtrip.py` (the `rt = make_runtime()`
  binding was never used; `@protect` resolves the runtime from the
  context, not the local) plus one I001 (DIGEST_PREFIX is upper-case
  so the import sort moves it ahead of BusinessImpact). ruff CI
  runs over `src/` AND `tests/`, so all three would have failed a
  build.

## Verification

| Check | Result |
|---|---|
| `ruff check src tests` | All checks passed |
| `mypy src/nullrun` | Success: no issues found in 37 source files |
| `pytest -q` | **1694 passed, 1 skipped** in 109.64s (vs baseline 1597 / 1 — **+97 new tests**) |
| Scratch diff | clean (no `dist_local/`, no `*.defect*`) |
| `nullrun.__version__` | `0.21.0` |
| Wire-format | additive on `/gate` (carries `business_impact`, `tool_class`, `mcp_annotations` when set; absent means "unknown", not "false"); non-additive on `/track` (the v3 single path now raises typed enforcement rejections where it dropped them — caller-observable) |

## Commits included

```
4bbb54e style: clear the three lint errors this branch's own test files carry
d11eb78 test(protect): pin the unbuildable-envelope degradation
88ddec8 fix(protect): build the tool_call envelope before the /gate pre-flight
015407b feat(gate): send the context envelope at /gate instead of a sentinel
e3176d9 feat(sdk): carry one BusinessImpact envelope per logical action
4320ac0 feat(sdk): restore the tool_call BusinessImpact constructor
6108308 docs(kill): stop claiming the kill signal is BaseException-only
b64dc8a fix(mcp): forward tool class and annotations to /gate
ea7c9ee fix(track): propagate enforcement rejections from the v3 /track path
9ccf168 fix(sdk): read the attribute the WS connection actually has
76efa7b fix(sdk): type the /gate pre-flight refusal instead of assuming budget
4ec5460 fix(sdk): register the two backend budget codes in the SDK catalog
```

(The version-bump commit and the `uv.lock` stamp `0.20.0 → 0.21.0`
are folded into this release commit.)
@codecov

codecov Bot commented Oct 2, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.23810% with 6 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/nullrun/business_impact.py 96.66% 1 Missing and 1 partial ⚠️
src/nullrun/decorators.py 84.61% 1 Missing and 1 partial ⚠️
src/nullrun/runtime.py 94.73% 0 Missing and 2 partials ⚠️

📢 Thoughts on this report? Let us know!

@maltsev-dev
maltsev-dev merged commit 7f057cf into master Oct 2, 2026
5 checks passed
@maltsev-dev
maltsev-dev deleted the release/0.21.0 branch October 2, 2026 08:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant