Skip to content

research: trial ecosystem responsibility routing #444

Description

@Jammy2211

Overview

Evaluate the routing convention proposed by research #440/#441 after terminology adoption #442/#443. The human asked "ok do it" in response to the recommendation to trial agent routing across visualization, profiling and inference; this task executes that approved research step.

Plan

  • Fix six hypothetical cases, identical evidence, expected routes and scoring before running evaluators.
  • Compare fresh-context responses under current guidance versus current guidance plus the proposed checklist, using the same model and effort.
  • Score decisions and proposed context selections against the predeclared key; retain no-change as a valid outcome.
  • Write a cited report and retain bounded raw/protocol evidence; independently review the conclusions and validate docs/JSON.
Detailed plan

Primary and only modified repo: PyAutoBrain, clean on main; no active claim conflict.
Branch: feature/ecosystem-routing-trial.
Worktree: /home/jammy/Code/PyAutoLabs/.worktrees/ecosystem-routing-trial/PyAutoBrain.

Write docs/research/ecosystem_routing_trial.md and bounded JSON/text evidence under
research/ecosystem_routing_trial/. Six cases (two/domain) test visualization
producer/organ ownership, profiling comparability/library changes, and inference
science/code routes. Fix expected answers, scoring dimensions and confounds before
evaluation; keep answer keys hidden from evaluators. Preserve exact requests,
responses, source commits and resolved model/settings. Assess producer, decision
owner, change owner, existing door, boundary and suggested minimal context.
Do not equate selected context or response length with observed retrieval savings.
This is a one-pass illustrative comparison, not a statistically powered benchmark.

Use read-only isolated evaluation sessions; no scientific execution or repo writes
by evaluators. Baseline uses current guidance; treatment adds only the proposed
routing convention. Keep same source evidence, case order and output contract;
report order/one-sample/hypothetical limits. Independent review checks raw evidence
and conclusions; run Sphinx/link/JSON/whitespace checks. No new framework, routing
policy, metadata, organs or unrelated tasks. Any follow-up ideas are proposed
before writing the ideas registry. Heart entry feed STALE; shipping gate separate.

Original prompt

Trial ecosystem role routing against current guidance

Type: research
Target: @PyAutoBrain
Repos:

  • PyAutoBrain
    Difficulty: medium
    Autonomy: supervised
    Priority: normal
    Consequence: judge
    Filed: 2026-10-01
    Status: draft

Original user request (verbatim)

ok do it

Accepted recommendation and context

In response to the completed layer-design research and terminology adoption,
the assistant recommended: "Trial the agent-routing convention. Use visualization,
profiling and inference examples to check whether agents correctly distinguish
the evidence producer, decision owner and repository needing a change. This is
the next step I recommend." The user authorized that next step above.

Research #440 / PR #441 proposed the convention; adoption #442 / PR #443
established responsibility terminology but did not change routing machinery.

Plan

Run a bounded exploratory comparison of six hypothetical routing cases (two per
domain), using fresh isolated evaluation contexts and identical case/evidence
packets. Baseline is current guidance after terminology adoption; treatment adds
the proposed evidence/producer/decision/change-owner checklist. Use the same
available model and effort for both; record resolved model and settings. Do not
present model identity or a hypothesis as a preferred answer to evaluators.

Before evaluation, fix the protocol, cases, expected decisions and scoring.
Include project versus organ visualization defects, environment versus library
profiling drift, and scientific inference follow-up versus a code defect. Missing
information must permit a conditional route, not an invented owner or verdict.
Do not run scientific jobs or use the trial to change scientific conclusions.

Score evidence attribution, decision owner, change owner, existing action door,
approval/record boundary and proposed minimal context. Distinguish predicted
context choices from observed tool reads; do not claim token/cost savings from
answer length. Report per-case paired outcomes, not statistical significance.
Retain no-change as a valid outcome; recommend a larger/real-world trial only if
justified by the results. Do not alter runtime routing, role schemas or organs.

Deliverables

  • docs/research/ecosystem_routing_trial.md: protocol, outcomes, counterexamples,
    limits, and recommendation on whether to adopt the checklist, revise it, or
    retain current guidance.
  • research/ecosystem_routing_trial/: bounded versioned protocol/cases, evaluator
    responses and assessment, with enough source/model/hash provenance to inspect
    the comparison. No new evaluation framework or reusable harness.
  • Independent review of conclusions against raw evidence; Sphinx/link/JSON checks.

All outputs stay in Brain. Other repositories are read-only evidence, with no
new claims. The separately filed inference documentation and profiling-board
registration tasks are not silently folded in. Candidate follow-up work is
proposed before being appended to Mind ideas, per the research skill.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions