Overview
Evaluate the routing convention proposed by research #440/#441 after terminology adoption #442/#443. The human asked "ok do it" in response to the recommendation to trial agent routing across visualization, profiling and inference; this task executes that approved research step.
Plan
- Fix six hypothetical cases, identical evidence, expected routes and scoring before running evaluators.
- Compare fresh-context responses under current guidance versus current guidance plus the proposed checklist, using the same model and effort.
- Score decisions and proposed context selections against the predeclared key; retain no-change as a valid outcome.
- Write a cited report and retain bounded raw/protocol evidence; independently review the conclusions and validate docs/JSON.
Detailed plan
Primary and only modified repo: PyAutoBrain, clean on main; no active claim conflict.
Branch: feature/ecosystem-routing-trial.
Worktree: /home/jammy/Code/PyAutoLabs/.worktrees/ecosystem-routing-trial/PyAutoBrain.
Write docs/research/ecosystem_routing_trial.md and bounded JSON/text evidence under
research/ecosystem_routing_trial/. Six cases (two/domain) test visualization
producer/organ ownership, profiling comparability/library changes, and inference
science/code routes. Fix expected answers, scoring dimensions and confounds before
evaluation; keep answer keys hidden from evaluators. Preserve exact requests,
responses, source commits and resolved model/settings. Assess producer, decision
owner, change owner, existing door, boundary and suggested minimal context.
Do not equate selected context or response length with observed retrieval savings.
This is a one-pass illustrative comparison, not a statistically powered benchmark.
Use read-only isolated evaluation sessions; no scientific execution or repo writes
by evaluators. Baseline uses current guidance; treatment adds only the proposed
routing convention. Keep same source evidence, case order and output contract;
report order/one-sample/hypothetical limits. Independent review checks raw evidence
and conclusions; run Sphinx/link/JSON/whitespace checks. No new framework, routing
policy, metadata, organs or unrelated tasks. Any follow-up ideas are proposed
before writing the ideas registry. Heart entry feed STALE; shipping gate separate.
Original prompt
Trial ecosystem role routing against current guidance
Type: research
Target: @PyAutoBrain
Repos:
- PyAutoBrain
Difficulty: medium
Autonomy: supervised
Priority: normal
Consequence: judge
Filed: 2026-10-01
Status: draft
Original user request (verbatim)
ok do it
Accepted recommendation and context
In response to the completed layer-design research and terminology adoption,
the assistant recommended: "Trial the agent-routing convention. Use visualization,
profiling and inference examples to check whether agents correctly distinguish
the evidence producer, decision owner and repository needing a change. This is
the next step I recommend." The user authorized that next step above.
Research #440 / PR #441 proposed the convention; adoption #442 / PR #443
established responsibility terminology but did not change routing machinery.
Plan
Run a bounded exploratory comparison of six hypothetical routing cases (two per
domain), using fresh isolated evaluation contexts and identical case/evidence
packets. Baseline is current guidance after terminology adoption; treatment adds
the proposed evidence/producer/decision/change-owner checklist. Use the same
available model and effort for both; record resolved model and settings. Do not
present model identity or a hypothesis as a preferred answer to evaluators.
Before evaluation, fix the protocol, cases, expected decisions and scoring.
Include project versus organ visualization defects, environment versus library
profiling drift, and scientific inference follow-up versus a code defect. Missing
information must permit a conditional route, not an invented owner or verdict.
Do not run scientific jobs or use the trial to change scientific conclusions.
Score evidence attribution, decision owner, change owner, existing action door,
approval/record boundary and proposed minimal context. Distinguish predicted
context choices from observed tool reads; do not claim token/cost savings from
answer length. Report per-case paired outcomes, not statistical significance.
Retain no-change as a valid outcome; recommend a larger/real-world trial only if
justified by the results. Do not alter runtime routing, role schemas or organs.
Deliverables
docs/research/ecosystem_routing_trial.md: protocol, outcomes, counterexamples,
limits, and recommendation on whether to adopt the checklist, revise it, or
retain current guidance.
research/ecosystem_routing_trial/: bounded versioned protocol/cases, evaluator
responses and assessment, with enough source/model/hash provenance to inspect
the comparison. No new evaluation framework or reusable harness.
- Independent review of conclusions against raw evidence; Sphinx/link/JSON checks.
All outputs stay in Brain. Other repositories are read-only evidence, with no
new claims. The separately filed inference documentation and profiling-board
registration tasks are not silently folded in. Candidate follow-up work is
proposed before being appended to Mind ideas, per the research skill.
Overview
Evaluate the routing convention proposed by research #440/#441 after terminology adoption #442/#443. The human asked "ok do it" in response to the recommendation to trial agent routing across visualization, profiling and inference; this task executes that approved research step.
Plan
Detailed plan
Primary and only modified repo: PyAutoBrain, clean on main; no active claim conflict.
Branch:
feature/ecosystem-routing-trial.Worktree:
/home/jammy/Code/PyAutoLabs/.worktrees/ecosystem-routing-trial/PyAutoBrain.Write
docs/research/ecosystem_routing_trial.mdand bounded JSON/text evidence underresearch/ecosystem_routing_trial/. Six cases (two/domain) test visualizationproducer/organ ownership, profiling comparability/library changes, and inference
science/code routes. Fix expected answers, scoring dimensions and confounds before
evaluation; keep answer keys hidden from evaluators. Preserve exact requests,
responses, source commits and resolved model/settings. Assess producer, decision
owner, change owner, existing door, boundary and suggested minimal context.
Do not equate selected context or response length with observed retrieval savings.
This is a one-pass illustrative comparison, not a statistically powered benchmark.
Use read-only isolated evaluation sessions; no scientific execution or repo writes
by evaluators. Baseline uses current guidance; treatment adds only the proposed
routing convention. Keep same source evidence, case order and output contract;
report order/one-sample/hypothetical limits. Independent review checks raw evidence
and conclusions; run Sphinx/link/JSON/whitespace checks. No new framework, routing
policy, metadata, organs or unrelated tasks. Any follow-up ideas are proposed
before writing the ideas registry. Heart entry feed STALE; shipping gate separate.
Original prompt
Trial ecosystem role routing against current guidance
Type: research
Target: @PyAutoBrain
Repos:
Difficulty: medium
Autonomy: supervised
Priority: normal
Consequence: judge
Filed: 2026-10-01
Status: draft
Original user request (verbatim)
ok do it
Accepted recommendation and context
In response to the completed layer-design research and terminology adoption,
the assistant recommended: "Trial the agent-routing convention. Use visualization,
profiling and inference examples to check whether agents correctly distinguish
the evidence producer, decision owner and repository needing a change. This is
the next step I recommend." The user authorized that next step above.
Research #440 / PR #441 proposed the convention; adoption #442 / PR #443
established responsibility terminology but did not change routing machinery.
Plan
Run a bounded exploratory comparison of six hypothetical routing cases (two per
domain), using fresh isolated evaluation contexts and identical case/evidence
packets. Baseline is current guidance after terminology adoption; treatment adds
the proposed evidence/producer/decision/change-owner checklist. Use the same
available model and effort for both; record resolved model and settings. Do not
present model identity or a hypothesis as a preferred answer to evaluators.
Before evaluation, fix the protocol, cases, expected decisions and scoring.
Include project versus organ visualization defects, environment versus library
profiling drift, and scientific inference follow-up versus a code defect. Missing
information must permit a conditional route, not an invented owner or verdict.
Do not run scientific jobs or use the trial to change scientific conclusions.
Score evidence attribution, decision owner, change owner, existing action door,
approval/record boundary and proposed minimal context. Distinguish predicted
context choices from observed tool reads; do not claim token/cost savings from
answer length. Report per-case paired outcomes, not statistical significance.
Retain no-change as a valid outcome; recommend a larger/real-world trial only if
justified by the results. Do not alter runtime routing, role schemas or organs.
Deliverables
docs/research/ecosystem_routing_trial.md: protocol, outcomes, counterexamples,limits, and recommendation on whether to adopt the checklist, revise it, or
retain current guidance.
research/ecosystem_routing_trial/: bounded versioned protocol/cases, evaluatorresponses and assessment, with enough source/model/hash provenance to inspect
the comparison. No new evaluation framework or reusable harness.
All outputs stay in Brain. Other repositories are read-only evidence, with no
new claims. The separately filed inference documentation and profiling-board
registration tasks are not silently folded in. Candidate follow-up work is
proposed before being appended to Mind ideas, per the research skill.