Repository navigation
perf: faster browser tasks, settling, hedged Jev calls and fewer rescues (openhuman#7000) - #77
Open
YellowSnnowmann wants to merge 33 commits into
Open
YellowSnnowmann wants to merge 33 commits into
YellowSnnowmann wants to merge 33 commits into
Conversation
A task's journal covered only its flows: the planner's call before the first flow, each rescue between flows, and a person's answer to a pause left gaps that read as unexplained time (38% of a live batch). The task controller now times them and hands plan, rescue, and resume events to a new FlowRunner::journal hook; the module's runner writes them into the task's file, or for PlanTask, which plans before a task exists, into a run of its own. Planner::plan_measured and Rescuer::guide_measured count the model calls each took, repairs of refused answers included. jev_journal gains --split, which reads whole tasks by their timestamps (each run restarts elapsed_ms) and splits wall time into planning, rescues, waits, and flows (Jev, settling, acting, reading), with per-call and per-decision latency and what the slowest call adds to each round; and --compare, the medians of two sets of tasks side by side. No change to what a task does: the hook does nothing by default, and nothing when the journal is off.
Picks up 5efb633 on the local agent-browser branch tinycomputer/network-quiet (not yet pushed to tinyhumansai/agent-browser): waitforloadstate gains networkquiet, which counts its 500 ms of quiet from the start.
Each is a setting so it can be compared against the current behaviour on the same tasks before it becomes the default (openhuman#7000, Stage 1). - browser.settle = "prompt" (TINYCOMPUTER_BROWSER_SETTLE): after an action, wait for the network to go quiet counted from the start (networkquiet), then only while the page still changes: no DOM change for 120 ms and no finite CSS animation running, over two frames, at most 400 ms. An idle page is read again after about 0.6 s instead of 1.6 s; a busy one still waits for its requests. - planner.plan_reasoning = "off" (TINYCOMPUTER_PLAN_REASONING): the planner asks its model not to reason first, with OpenRouter's reasoning.enabled=false. Live, a plan took about 3.5 s instead of 16 to 20 s; reasoning_effort and thinking were ignored on that route. - browser.prelaunch (TINYCOMPUTER_BROWSER_PRELAUNCH): StartTask with a plain-language task opens a browser-only task's browser while it is planned (FlowRunner::prepare, BrowserSurface::open), and lets it go if planning fails. task_live's TASK_PLAN=in-task plans inside StartTask, as OpenHuman does, so this can be measured.
planner.plan_reasoning = "off" cut a plan from 16-20 s to about 3.5 s, but in the Stage 1 A/B every extra rescue traced back to a plan written without reasoning: card fields filled before the stop, "press Enter" where only a button searches, a location step that stalled, a verify step that failed. The time saved was lost again to rescues, so plans are drafted with reasoning again and the switch is removed.
browser.settle now defaults to prompt and browser.prelaunch to true. In 44 live runs over the two Stage 1 A/B batches, no failure traced back to either: Myntra's size showed selected straight after its click, and the failures were the irreversible-press gate, a store's own error overlay, and plans drafted without reasoning. Settling promptly cut the wait after an action by 0.2-0.7 s on three of four stores (Amazon's network never goes quiet, so it still waits about 2.3 s), and opening the browser while the plan is drafted cut the first step's launch from 2.8 s to 0.6 s. Both stay switchable: "settle": "steady" and "prelaunch": false restore the old behaviour, and task_live takes TINYCOMPUTER_BROWSER_SETTLE=steady and TINYCOMPUTER_BROWSER_PRELAUNCH=0.
Sight hands a page to the accessibility tree when it sees a control inside a shadow root, which a CSS selector from the page cannot address. It only counted a shadow root whose host itself showed. On Lenskart the consent banner's host is laid out as `display: contents`, so it has no box, sight skipped it, and the banner's "Allow Selection" and "Allow all" buttons never reached Jev while the banner lay over "Add To Cart": every press came back covered, and all four runs failed after five rescues. A shadow root now counts once its host or any of its controls shows. The live fixture test reproduces the banner (sight read only "Add To Cart" before the fix), checks a block host whose banner is fixed and draws no box either, keeps a hidden banner out, and checks that the tree read instead offers the banner's buttons.
After text goes into a place box (an address, a city, a pickup), the flow looked twice more for suggestions the page lists late, each look a fixed 500 ms pause plus a full settle. A plain form box lists none, so every address and city field on BlazeDemo's passenger form paid 2.3 s (4.2 s with steady settling) for a list that never comes, four waits a run. Surface::await_change waits for the surface to change by itself and says whether it did. The browser watches the page with a MutationObserver that ends at the first change of its own (never sight's data-tc- marks), or after LATE_LOOK_MS = 1000 ms with none; a surface that cannot watch, the desktop's, pauses as a Wait did and says it may have changed. The flow looks again after each wait, as before, but a wait that saw the page stay still is not settled, notes "nothing changed", and ends the looking: a still page lists nothing more. Rows that arrive late are looked at as soon as they are drawn. The simulated ride form now draws its rows a number of waits late; the new flow tests pick rows that come one and two waits late, and wait once, not twice, beside a box that lists nothing.
Asked for the price of one packet after raising its count to 2, Zepto runs read the cart line's ₹40, the total for both: the cart shows no price for one, and the product page's ₹20 lay behind the cart drawer. The flow guide, which the planner and the rescuer both read, now says to read the price of one before the count is raised.
One HTTP 502 from Tiny Humans' gateway failed a whole BlazeDemo task at its plan, 3 s in: a hosted model call that failed was never tried again. It now is, up to 4 times, 1, 2, then 4 seconds apart, when the failure can pass: a server error, a rate limit that is not a spending cap, or a dropped connection (tinyinference-llm's own classification). A refused key, a rejected request, and the test guard's refusal are not retried.
In the 7 October runs, 21 of 36,459 Jev calls needed a retry: one framing of a burst stalled while its siblings answered in about 0.6 s, until the gateway gave up after ~10 s (12 s with the retry) or nothing came back before the client's 30 s timeout (31.6 s). Every decision waits for all of its framings, so about one run in four lost 12-32 s to one of them. A framing that has not answered after HEDGE_AFTER (2.5 s; 3.5 s for a request of 32 KB or more, whose p99 was 3.1 s) is now sent once more, and whichever copy answers first counts; a copy that fails gives way to the other. Fewer than 0.3% of calls ran that long otherwise, so the copies cost little. The answer that counts carries the extra attempt, and the journal records a `hedge` event. Every framing goes through it, the evidence gate's included. A Jev attempt now gives up after 10 s unless jev.timeout_ms says otherwise (the slowest answer seen took 8.6 s), so a request both copies of which stall is retried after 10 s rather than 30 s.
The networkquiet wait now counts only requests that can change what the page shows (the page's own document, scripts, stylesheets, XHR and fetch data), so analytics pings and other frames' documents that never report finishing no longer hold every settle to its cap.
On Amazon every action that opened a page sent 100+ requests for over 2 s, so prompt settling always ran its network wait to the 2 s cap, then its 400 ms stillness cap, about 2.4 s an action. What the task needed was on screen by then for a long time: the results after 0.87-1.04 s, a product's title and Add to Cart after 0.96-1.24 s, the cart's subtotal after 0.66-0.81 s (six measured page loads). Settle::Prompt now waits at most QUIET_MS = 1 s for the requests that change the page, then for the page to stop changing as before, so such a page is read after about 1.4 s. Steady settling keeps its 2 s.
In the live check after the rule first went in, 3 of 4 plans read the price of one before raising the count, but one Zepto plan still read it in the cart and got ₹40, the line for two. The rule now also sits in the `read` row of the step table, where a planner looks when it writes a read.
Handing the whole page to the tree whenever a shadow root showed controls made Lenskart readable but worse to work with. While its consent banner was up, the tree read the search box unnamed, so Enter did not search; it cut the page at its element budget; and it read an image carousel's dots as sizes. The banner reached Jev in 1,330 calls of one run and was still never closed. Sight now keeps reading the page. For a shadow root that shows controls (its host's box or any control's), it marks the host and names the layer the shadow content draws, from its fixed or dialog element's label, heading, or first words: `popover "We value your privacy"`. The surface reads the tree under that host alone and adds its controls and text after everything sight read, under that label, so the digest sees a layer in front and the attention pass a privacy card with "Allow Selection" to close it with. A tree snapshot's refs last until the next snapshot, so with two such shadow roots, or when the subtree cannot be read, the tree reads the whole page as before.
A press the page refused as covered was answered with Escape and one
more try. Escape leaves a consent banner where it is: on Lenskart every
press of "Add To Cart" and "Frame Size" under the banner was refused
until the run failed.
When a layer in front holds a control that closes it, the least committal
one ("Allow Selection" before "Allow all", as the attention pass ranks
them) is now pressed instead, as `click (uncover)`, and the same target
tried again; otherwise Escape, as before. The simulator gains a consent
banner Escape does not close.
On a slow evening the Tiny Humans route answered calls in up to 3.4 s (median 0.91 s, against 0.72 s earlier in the day), and a copy sent at 2.5 s lost the race to its original 15 times in 16: 5% more calls for nothing. A framing now gets its copy after 4 s, or 5 s for a request of 32 KB or more (p99.9 3.9 s), past what a slow but live answer takes; a stuck one still answers in about 5 s rather than 12-32 s.
A `do` step whose last three actions changed nothing failed outright, even when the page had already done its work: "press Enter to search" pressed Enter three times over results a live search had listed as the query was typed. On 7 October, 25 of 189 rescues found such a step's work already done, each ~18 s after the failure. Such a step now asks one question first: does the screen already show the result the step is meant to bring about? When it clearly does (DONE), the step ends AlreadyDone; otherwise it fails with the same note as before, so a rescue still skips rather than retries it.
A task's report lists the elements it grounded (`learned`), meant to be passed back as StartTask's `memory` so a later run confirms a remembered element with one yes or no rather than searching for it. Neither task_live nor OpenHuman passed any. TASK_MEMORY names a JSON file of grounding hints: the task starts with them, and what it learns is saved back, a newer hint replacing an older one for the same element. Unset, every run starts fresh, as benchmark batches should.
…w root Reading one shadow root's controls beside sight snapshots the tree under its host. On Lenskart that snapshot failed (agent-browser described the host's subtree without piercing its shadow root, so no accessibility node was found), and the surface fell back to reading the whole page through the tree, as before the change. agent-browser d8e93d8 describes the subtree through shadow roots. The live fixture now puts a stylesheet beside the banner in its shadow root, as the live banner had.
agent-browser now keeps each page session's page-changing requests that have not finished, and a networkquiet wait starts with those sent in the last 3 s: a click's own navigation or fetch, which starts before the wait can subscribe, is waited for, up to QUIET_MS, instead of the page being read as it was before the action. The bump also brings the networkquiet docs, help and MCP entries, and the subtree doc comment fix. Settle's docs say so, and give their waits as values: rustdoc refuses links from a public enum to private constants, which failed the docs job.
The prompt settle's stillness script and await_change's watch were sent with no deadline of their own, right after the actions that replace pages, and an evaluate sent while a page is replaced waited out the browser's 30 s live. Each call now has its own cap plus 500 ms (watch::deadline): a quiet wait given up on still lets the page be watched, and a watch given up on says the page may have changed. Both watches now observe the page's open shadow roots as well as the document, so a suggestion list a web component draws counts as a change. The stillness script stops asking for frames once it has resolved, and await_change with no page open opens none. The scripts move from operations.rs to their own watch.rs, and the test engine can stall a command to prove the deadlines hold.
With one shadow root showing controls, the tree reads its host's whole subtree beside sight: the host itself and the light-DOM children the page puts in its slots, which sight had already read. Each was then offered twice, as seen:N and eN, and a host that slots the whole page (an app shell) would have listed the page twice. Sight now leaves the first such host and everything under it to the tree. Only the first host is labelled (with two, the tree reads the page anyway); a layer that is not shown, such as a hidden dialog template, names none; and a banner's aria-labelledby resolves inside its shadow root. A reading the tree takes over no longer reports sight's denoising counts. A live fixture test covers the slotted button, the hidden dialog, and the label.
A browser-only task whose plan asks for values published needs_input with the browser it opened while planning still open, holding one of the browser's sessions for as long as the person took, or for good if the host never answered. It is now released first; the run that follows opens one again. A task cancelled while its early open had not begun yet could leak a session: release found none to close, and the open then launched one nobody owned. A closed BrowserSurface now refuses an early open, checked under the session lock, so either the open sees the close or the close waits for the open and closes what it made.
Sage's calls take 5-6 s as a rule, past any hedge delay: almost every framing would have been copied, adding half again to its calls and cost for answers that rarely came first. Sage now gets no copy. Many framings outliving their delay together means a slow or failing gateway, where the client may be waiting out its own retry delay; a copy of each only loaded it more. A runtime now has at most HEDGE_COPIES (2) copies in flight; past that, a framing waits for its own answer. The hedge event's won now names the answer that counted (first, copy, or neither), not the copy that finished first. Hedging moves from decide.rs into its own module, and the harness doc's settle row says how the browser settles now.
When a press was refused as covered, front_closer took the least committal closer of the first layer in front, asked with no intent and never told the target. A toast over a size popover's rows could get the popover's own close pressed, and a chat bubble over a consent banner's Accept all the banner's Reject all: either closes the layer the step works in, and the retried press then fails. It now skips the layer the target sits in, the target itself, and any layer the step's intent names, as the attention pass does. The covered press moves into act/uncover.rs, and its docs and the attention module's say what it presses; decision-loops.md is tightened around it so the file does not grow.
Whether a step is done is the completion loop's question, and every other completion question in the do loop is gated on it; the stalled-step check asked holds() even when a run turned that loop off. Such a run now fails a stalled step outright, as before the check. decision-loops.md, decision-thresholds.md (STALL_TURNS and DONE) and jev-questions.md still described a stall as an outright failure, and now say what happens instead.
The simulator now keeps a trail of its settle and await_change calls. The late-suggestion tests check that a wait that saw a change is settled like any action, and one that saw a still page is not.
The plan, rescue and resume events were built for every task, journal on or off, and journal_event drew a run id even when off: the journal's own rule is to build nothing then. FlowRunner::journal and journal_event now take a closure, called only when an event is written, and the module's runner and JevRuntime::journal_event return before anything is built when the journal is off (JevRuntime::journaling). A resume was journaled before ContinueTask's answer was checked, so a refused answer, or one still missing values, recorded a wait the task had not left, and --split counted it twice. It is journaled once the task has left the wait. A plan's wall_ms is timed on the planner alone: a browser slower to open than the plan was counted as planning. ModelUse's calls are documented as excluding a hosted call's own retries.
TASK_MEMORY's own example, target/task-live/memory/amazon.json, names a folder nothing made: the whole task ran, then saving what it learned failed and the run ended with an error before its pass or fail line. The folder is now made when missing. TASK_PLAN other than in-task, and TINYCOMPUTER_BROWSER_PRELAUNCH other than 0 or 1, were silently ignored; they are now refused. A FLOW_FILE run with TASK_PLAN=in-task no longer saves the given flow as a plan drafted inside the task. What a run keeps beside its report moves into its own module, so main.rs stays under 400 lines.
split.rs had grown to 537 lines. Building a split stays in split/mod.rs; rendering it for a terminal moves to split/render.rs, and reading an event's time and merging spans of it to split/time.rs. Nothing changes in what it prints.
The docs said task_live writes its journal under TASK_OUT/journal, and showed try-* folders. Neither task_live nor tasks/run sets the journal's folder: a run journals there only when started with TINYCOMPUTER_JEV_JOURNAL=$TASK_OUT/journal, and try-* was one private script's naming. The docs, jev_journal's help and runs.rs now say so.
Docs still said the browser settles on network idle (decision-loops.md, the-do-loop.md), that a place box waits a fixed beat between its late looks (filling-forms.md), and that a Jev attempt's timeout defaults to the client's (JevConfig::timeout_ms, jev-runtime.md); it is the module's 10 s. Three paragraphs of decision-loops.md are reflowed so the file does not grow past its length on main.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Comment |
YellowSnnowmann
marked this pull request as ready for review
October 7, 2026 14:43
Tiny Sweeper review
|
A place box was looked at again only while it listed no new row at all.
Live on Uber, the pickup box first listed rows of its own ("Allow
location access", "Search in a different city"), and the prompt settle,
which reads the page as soon as it goes still, looked before the matches
were fetched: no row named the place, Jev rightly chose none, and the
pickup was never set, so no ride options showed. The slower settle had
hidden this by looking later.
A place box is now looked at again, as before up to LATE_LOOKS waits that
each end at the page's first change, until a row names the text typed.
The simulator's ride form can list such rows of its own while its matches
are pending, and the new test fails on the old condition.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This makes browser tasks faster without losing reliability, for tinyhumansai/openhuman#7000. Each change was measured on live runs before it went in. The branch then had a final review, whose fixes are 13 commits, and a live run of ten utility sites on the final code, whose one regression is fixed in the last commit.
Measuring (Stage 0)
jev_journal --splitbreaks a task's time down into planning, Jev, settling, actions and rescues.jev_journal --comparesets two batches side by side.Less waiting (Stage 1, now default)
browser.settle = "prompt": after an action, the page is read once the requests that change it are done and it stops changing, about 0.6 s on an idle page instead of 1.6 s.browser.prelaunch = true: a browser-only task's browser opens while its plan is drafted, which takes the first launch from 2.8 s to 0.6 s."steady"andfalserestore the old behaviour.Shorter Jev waits (Stage 2)
jev.timeout_mssays otherwise.networkquiet). Amazon waited ~2.4 s after every page-opening action: 100+ requests over 2 s, some of which never report finishing to the page's session. Its content shows after 0.7–1.2 s, and settling now takes 0.67–1.47 s a page load.Fixes from the live checks
Surface::await_change). A still page ends the waiting: BlazeDemo's address and city boxes went from four waits of 1.2–2.1 s to two of 1.0 s.display: contentshost that sight never saw.Fewer rescues (Stage 3)
dostep whose last three actions changed nothing first asks whether the screen already shows its result, and endsAlreadyDonewhen it clearly does. 25 of 189 rescues on 7 October were such steps, ~18 s each.TASK_MEMORY); it is off unless set.Fixes from the final review
jev_journal'ssplit.rsandtask_live/main.rsare split under 400 lines, and stale docs are brought up to date.Fix from the live run of the utility sites
LATE_LOOKSwaits that each end at the page's first change, until a row names the text typed.Checked and skipped, with data
Related issue
Refs tinyhumansai/openhuman#7000.
Depends on tinyhumansai/agent-browser#2. The
vendor/agent-browsergitlink pins its head,ebf10fb, which CI fetches through that pull request. If it is squash-merged, the gitlink must move to the merged commit before this one merges, so this PR is a draft until then.API or behavior changes
browser.settle(promptdefault, orsteady) andbrowser.prelaunch(truedefault).jev.timeout_msnow defaults to 10 s, from the client's 30 s.AlreadyDone;Surfacetrait: a newawait_changemethod with a default (pause and assume a change), so existing implementations compile unchanged.FlowRunner::journaltakes its fields as a closure, built only when written;JevRuntime::journal_eventlikewise, andJevRuntime::journalingsays whether the journal is on.plan,rescue,resumeandhedgeevents; awaitaction notes "nothing changed".jev_journal:--splitand--compare.TASK_PLAN=in-task,TINYCOMPUTER_BROWSER_SETTLE,TINYCOMPUTER_BROWSER_PRELAUNCH,TASK_MEMORY.None of this breaks the wire contract.
Validation
Commands run (Rust 1.99, macOS), with their outcomes:
cargo fmt --all -- --check: clean.cargo clippy --workspace --exclude tinycomputer-accessibility --all-targets --all-features -- -D warnings: clean.tinycomputer-accessibilityalready fails clippy 1.99 on macOS onmain(its macOS-only files), and this branch does not touch it.cargo build --workspace --exclude tinycomputer-accessibility --all-targets --all-features: ok (built by the test and coverage runs).cargo test --workspace --exclude tinycomputer-accessibility --all-features: 1,136 passed, 0 failed; after the last commit,cargo test -p tinycomputer-engine --all-features: 420 passed, and its clippy is clean.RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace --exclude tinycomputer-accessibility: clean.cargo llvm-covon core, bus, browser, engine and the module crate, with the CI gate's per-file filter: all 181 files at or above 90%, every changed source file at 100%.TINYCOMPUTER_LIVE_BROWSER=1, real Chrome): all 11 sight live tests pass, including the shadow-root banner and its slotted content.Live A/B batches (
task_live, planned in-task as OpenHuman does):piercefix addressesThe final review's fixes ran in the live run of the utility sites. The last commit's place-box fix is covered by a simulator test that fails without it; Uber has not been run again with it. The open Stage 3 failures are on this branch's list of follow-ups, not caused by it: BlazeDemo's
stop_beforenot finding "Purchase Flight", Lenskart's consent banner asking for a second choice, and Blinkit adding extra items.Tests
won;AlreadyDone;settle/prelaunchkeys;await_changedelegation;TASK_MEMORYmerging and its folder, and task_live's switch values;display: contentshost; a host whose slotted button, hidden dialog andaria-labelledbyare read once and right.Documentation
HEDGE_AFTER,HEDGE_COPIES,LATE_LOOK_MSandLATE_LOOKS, and the stall'sSTALL_TURNSandDONE.TASK_OUT/journal.decision-loops.mdwas already 509 lines onmain; this branch leaves it at 508.Checklist
#[allow(...)],#[ignore], or relaxed lints.envcontents in the diff or the description