Skip to content

feat: browser agent fixes from live tests on ten utility sites - #76

Merged
senamakel merged 48 commits into
tinyhumansai:mainfrom
YellowSnnowmann:fix/task-live-fast-hosted-model
Oct 7, 2026
Merged

senamakel merged 48 commits into
tinyhumansai:mainfrom
YellowSnnowmann:fix/task-live-fast-hosted-model

Conversation

@YellowSnnowmann

@YellowSnnowmann YellowSnnowmann commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

These fixes come from live tests of the browser agent on real Indian utility sites: Uber, Ola, Amazon, Myntra, BookMyShow, BigBasket (round 1), then Flipkart, Lenskart, Blinkit, and Zepto with the same sites again (round 2). Each fix is generic: no hostnames, site labels, or per-site steps anywhere in the engine, sight, or the planner guide.

After round 2, Amazon, Myntra, Zepto, Flipkart, Lenskart, and Blinkit run end to end to the cart or the payment checkpoint, and Uber reaches its final "Request" button after the person signs in. Round 3 (cab booking, a grocery basket, cinema seats) takes Ola to its "Confirm & Book" gate, BigBasket to its OTP login with the right item ×2 in the basket, and BookMyShow through date, format, show time, and seat count to two adjacent seats chosen in the cheapest section ("Pay ₹400").

The whole PR was then reviewed end to end (section "Review" below): four reviewers read the full diff, every finding was checked against the code, and each real one was fixed with a test that fails without it, including three blockers (two ways a secret could leak, and a dialog that stayed "the task's own" for the rest of the run). No code path names a host, a site, or one site's labels.

Forms and suggestions

  1. enter picks the suggestion an autocomplete box lists.
    • A location or city box keeps its text only once a listed match is chosen. Before, the next step pressed Escape on the open list and the box emptied.
    • Now enter looks again after typing and picks the new row that matches. If several rows match, Jev chooses among at most 12 of them, and it may answer that none fits.
    • Live follow-up: on a store's search box, the completions are other searches. A pick of "boat airdopes 141 anc" at 0.43 (a near tie with "none fits") changed the search. A suggestion is now pressed only at 0.5 or more (SUGGESTION_FLOOR), and a lone row is pressed without asking only when it reads exactly as typed.
  2. A panel that strings its suggestions together is never pressed.
    • A store's delivery-area popover was read as one button whose name held every row. It was the only new element mentioning the pincode, so it was pressed.
    • The press landed on the row at its middle and set a different area. A label that lists far more than the typed text is no longer a suggestion.
  3. Sight reads the rows inside a panel.
    • Any element with a tab stop or a pointer cursor became a button, and every element inside it was skipped. That swallowed the popover's rows, and a tab panel's rows the same way.
    • A region role (menu, list, grid, tab panel, dialog, landmark) no longer counts as a control just for its tab stop. One the page makes pressable itself, such as a carousel slide with a pointer cursor, still does. A button or link that holds a box to type in is now read as a panel.

Requests and retries

  1. A request too large for one Jev call is asked in parts.
    • Knockouts over long results pages carried 16 to 19 groups, 73 to 79 KB. The Tiny Humans gateway refused them with HTTP 502 every time, which ended the task (Amazon, BookMyShow).
    • Such a request is now cut by its questions into parts asked at once, each with the whole state. Jev evaluates each question on its own, so the answers merge back by id.
    • The cap goes from 100 KB to 48 KB. Across 7,068 live calls, 57 KB (23,600 tokens) was the largest that passed and 68 KB the smallest that was refused.
    • Live follow-up: vote::ballots took its question ids from the first answered framing only, so every part but the first lost its answers. Live, a search's "Go" won its group at 0.98 in the second part and was never pressed. The ballots now cover every question any framing asked.
  2. A provider's brief outage is ridden out.
    • The client retried a 5xx twice, 100 ms and 200 ms apart, so a task ended 2.5 s into a gateway blip.
    • When max_retries is not set, a Jev call now retries four times, waiting 1, 2, 4, then 8 s.

Waiting, reading, safety

  1. wait_for stops when the page says it found nothing.
    • A "No Results Found" page was checked ten times over (~50 s), three times in one task, and each rescue guessed at the query's wording.
    • Two checks in a row showing "no results", "0 results", "no products found" and the like now fail the step. The note names a phrase from our own list, never the page's text.
  2. Values read into declared variables are reported.
    • The planner declares each read's variable up front, empty. The task dropped every declared variable as the caller's input, so a requested brand, price, or bag total never came back.
    • A variable now counts as read when the run changed it.
  3. Login walls.
    • "Log in to see / view …", "log in or sign up", "please log in", "you must be logged in", and "login required" now pause as needs_human.
    • A header's bare "Log in" and "Sign up" still do not count.
  4. The reCAPTCHA badge is no longer a captcha.
    • Only a challenge counts now: "complete the reCAPTCHA", "select all images / squares", hCaptcha's "I am human".
  5. A store's and a ride app's last buttons are gated.
    • "Place your order", "Confirm order", "Complete order", and "Proceed to pay" count as payment.
    • "Confirm ride" and "Confirm pickup" need approval.
    • "Book" alone stays reversible, since on travel sites it opens the traveller form.

Planner guide

  1. An option group (size, colour, quantity) is a choose of its label, never a pick. A filter on which result to take ("skip Sponsored items") belongs in the pick, never a step of its own. A pick keeps the task's item and criterion ("the first boAt Airdopes 141"), never a stand-in such as "lowest price", which took another brand live.
  2. A condition names a state ("the sort shows Price: Low to High"), never a change ("the grid has updated after sorting"). A change cannot be judged on one screen, and it held a sorted list under the bar for ten checks.

task_live

  1. Its hosted roles run on openrouter/deepseek/deepseek-v4-flash, the managed default OpenHuman uses, instead of agentic-v1.
  2. TASK_INTERACTIVE=1 waits for the person at the terminal instead of ending at every pause:
    • approve or decline an action;
    • get past a login or captcha, then press Enter;
    • type a missing detail;
    • finish on the payment page.
  3. It prints what the task read, and writes records.json, however the task stopped.
  4. A pause's reason now ends with one full stop in its prompt.

Round 2: stores, ride apps, and a cinema

Commits 60e7879b to 428fb3be.

Sight (what the page reads as)

  1. A box is named by its own label, aria label, placeholder, or title before the words beside it, and a divider's "OR" never names one: a search box read as "Location not set", a location box as "OR".
  2. An unmarked picture button beside a bare count is "decrease" or "increase"; icon glyphs (an icon font, or a lone glyph in a font other than its parent's) are pictures, not names ("p", "!"); an icon's sprite symbol or file name names it ("#icon-cart" is "cart").
  3. A small element whose React props hold a click handler is a button: a store's "Add to cart" and "Buy now" were plain words.
  4. Repeated plain boxes (three or more siblings of one tag and class, each with a link or button and a word) are cards, numbered across rows so a grid of rows reads as one list.

Lists and picks

  1. Runs of three or more same-role links, options, radios, or buttons showing 20+ characters are lists too: a store's link cards and a ride app's options were never offered. A list whose cards show only bare markers (carousel dots, chips, page numbers) is not results.
  2. A card opens through a control that says so, then a named link (its title), then the first control: a card opened through its pack-size dropdown.
  3. pick by "first " walks the list in page order for the first item that meets it (it took the fifth result, another model).

Browser presses

  1. Links and forms aimed at a new tab, and a script's window.open(url) right after a press, stay in the session's tab.
  2. A click refused as covered is tried once more with the element centred and the pointer moved off (a product photo's hover zoom covered "Add to cart" twelve times); a content link whose press went nowhere is followed to its address.
  3. Pressing a "use my current location" control grants the session geolocation first, standing in for the browser's Allow bubble.
  4. A navigation that timed out counts as open when the page is drawn; one page reading is capped at 10 s (a look took a minute).

The do loop and dialogs

  1. A dialog or layer the run's own press opened is its next stage, across steps and rescues: never cleared by attention, obstacles, Escape, or undo; covered controls and its close control are not offered as moves; a failed step names its controls for the rescuer. A layer that typing opens (suggestions) is the field's.
  2. A control or key pressed three times in a step is not pressed again; pressing a named control strikes off its copies on other items for the step ("Add" on six products for "add 2 packets").
  3. Attention keeps a layer open when a control in front fits the step; "allow selection" is a consent closer; a stalled step says its work may be done; a control out of view counts as pressable.

Typing and suggestions

  1. A plain step that types ("enter 560001 into the pincode field") runs as enter; an enter that typed nothing fails instead of reporting done, first presses a control named by the slot's word, and never types two slots into one box (a drop typed over a pickup).
  2. A search box takes only a suggestion that is the same search; a place box looks twice more for late rows, and counts rows already showing that match or share most typed words.

Steps, safety, planner, rescuer

  1. wait_for holds on a settled screen judged at 0.65+ three times; more "found nothing" phrases; a date option matches a strip's "WED 07 OCT"; a read takes the text holding the value itself.
  2. OTP sign-in dialogs ("Login/ Sign up Using OTP") and "Sign up or Log in" pause as needs_human.
  3. Planner guide and rescue protocol: dialog headings are not answers; no confirm step after a picked choice; delivery place only when the task gives one or the page asks; add before raising a count; seats by a plain step; a day in a date strip is a choose; never a stop_before for an unasked login; keep the task's size and variant words.

Round 3: cab booking, a grocery basket, and cinema seats

Commits fec1cca0 to b054d7d5.

  1. A pick the page already shows selected is taken as it is (fec1cca0): Uber's cheapest car was selected by default, and pressing it again opened a fare breakdown over "Request".
  2. A − count + stepper in place of Add means the item is in the cart (a7a5306e, planner guide): BigBasket's basket held the milk ×2 while the run looped to add it.
  3. Before a covered retry, a focused text box lets go (6e6abc47), closing the suggestions it held open over the target; rescues never add an item twice.
  4. A place box takes its closest row when none names the place exactly (9dc611f3): Ola lists "MG Road Metro Station" as "MG Road, Shivaji Nagar". The rescue protocol learns the date strip and in-card time rules.
  5. Sight reads a table's rows and buttons the way a person does (18b2f821), all from BookMyShow's accessible seat table:
    • an unlabelled row is named by its control-free cells, row "01 Available" #1, so a seat's status reaches Jev (it had pressed a Handicapped seat);
    • a merged wrapper (td role="gridcell") points at the native button inside it: six presses at the cells' middles selected nothing while the buttons sat at their left edge;
    • aria-hidden cells that take no pointer events count as seen: the hit test passed through them, and every row in view lost its seat number.
  6. A rescue that reruns the failed step covers nothing after it (86337b98): one claimed covers: 1, and the show-time step was dropped.
  7. Lists (25331ade): a list of one-line items that are another list's cards split apart is dropped (Ola's ride cards also read as their nine lines, losing names and fares); a list not clearly chosen goes to the one Jev leaned to (0.41 against 0.05) before the longest.
  8. Items chosen together read as one source (a390a15e): "the selected seats" could take only one piece of text.
  9. A step that chooses several items keeps the copies rule off (94181464): "choose 2 adjacent seats" had its second seat's "Select" struck off as another item's copy; "add 2 packets of milk" keeps the rule.
  10. The task's own dialog is answered when nothing serves the step (f55b43c6): a date step left a format dialog's "2D" alone for four turns and spent a rescue. A control the dialog's own bar covers (a seat row under "Pay") is offered; what lies behind the dialog still is not.
  11. Planning keeps a choice's conditions, and presses a booking button first (b054d7d5): "the cheapest car" became "lowest fare" and took an auto.

Review: the whole PR, end to end

Commits 6f242057, 1fcfd171, 35e46635. Four reviewers read the full diff against main: the browser adapter and sight; the do loop, dialogs, and the Jev plumbing; step kinds and form filling; and core, rescue, tasks, and the runner. Every finding was checked against the code before it was fixed.

Hardcoding check. No hostname, site name, or one site's label appears in any code path. The word lists are generic UI vocabulary (login walls, "found nothing", place words, number words, role names, cookie-consent labels). Site names appear only in comments that describe a live failure and in test fixtures; the site-flavoured examples in the planner guide and rescue prompt were reworded.

Fixed, each with a test that fails with the fix switched off:

  1. Secrets (blockers). A plain step that types was substituted twice: a value read off a page that said ${card_number} was typed as the caller's card number. It is now read from the step as written. A variable a flow defines from a fact ("recipient": "${email}") was reported as a read, which put a secret's value in the task's records; only a variable a step writes is a read now.
  2. The task's own dialog (blocker). A dialog a step opened stayed "the task's own" for every later step, so a calendar left open was never cleared again. A step that works in such a dialog now hands it back at the next step; a scroll or the run's own housekeeping (clearing, dismissing, undoing) opens no dialog of the task's; opening an address forgets it. Answering such a dialog when nothing serves the step (45) now happens only on a browser task, among the dialog's own controls, and never presses a control that commits ("Yes", "Confirm", "Pay", "Book").
  3. Typing steps. "Enter" also means going into something ("enter Reader mode in Safari", a desktop regression), so with " in " it types only what is quoted, data, or into a box; a step that does more than type ("… and press Enter") stays a plain step for a rescue to split; "type in" is a verb; "enter the ${otp}" types the code, not "the 123456"; a quoted text ends at its quote; the split prefers the field that names a box.
  4. Presses. A control's copies are struck off only on the other items of its own list (a sticky bar's "Add to cart" is no copy), only once the next look shows the press changed the screen (a refused press acted on nothing), and an undo lifts only that press's copies. A scroll counts toward no press cap, so a long list scrolls more than three times.
  5. Form filling. The control named by a slot's word is pressed only once Jev agrees it shows that slot's box (OPENER_FLOOR 0.8), never on a shared word or the role word "link" alone ("Email us" shares "email"). A place box with no fitting row asks Jev once more for the row naming the place in other words, instead of pressing the row sharing most words (a default click). A search box offers only new rows, so a menu link reading as the query is not pressed; a place box whose own words say "Search for area…" takes a place.
  6. Picks. "first X" is a condition only with a condition word (rated, under, with, available…): "First AC" and "first class" are names, "first to depart" an order Jev judges. Round 3's guide rule put a choice's kind in by, where a measured ranking ignores it; the guide and rescue prompt now put it in from ("the car options" by "lowest fare"), where the ranking checks it.
  7. Login stops. Found live on Amazon in this review: a plan for "do not log in" stopped short of the cart in front of the header's "Hello, sign in". A stop_before that names only signing in now gates nothing (a login wall pauses for a person by itself); signing up or creating an account stays gated.
  8. Rescues. Guidance that ends by rerunning the failed step covers only the later steps its earlier steps do (the same step, or an enter filling one of their fields): round 3's rule set covers to 0 outright, which reran fields a rescue had already filled.
  9. Reads and waits. A value read before a rescue outranks the flow's empty declaration of it (${total} expanded to nothing after a rescue). "no products" and "0 products" (a header's empty cart) no longer read as a search that found nothing.
  10. Jev requests. A screen that fills most of a request is cut before splitting, so its questions are not asked one call each, and the budget is charged per part and framing.
  11. Browser presses. Keeping a press in the tab changes only the pressed link or form (it rewrote every _blank link on the page), opens a script's new window in place only for an address on the same site (an advert's window opened on a click no longer takes the tab), and returns no window (a page's w.close() closed the agent's tab). Location is granted only in a browser the module launched on a throwaway profile, never in a person's own browser or profile. A link press is no longer followed for an expand or collapse, a download, a link that controls the page, a page already leaving, or behind a dialog shown in front; a failed follow keeps the press's own reply. The drawn-page check compares the query too. A covered retry keeps the focus when the target is a row of the focused box's own list.
  12. Sight. Only an element that turns pointer events off itself counts as seen through a hit on what holds it, never a hit on the page's root: a modal library turns them off for the whole body while it hides the page behind its dialog. Only a press handler (onClick, onPress) makes a React element a control (a carousel's onMouseDown track became one big button), and such a control hides nothing pressable inside it. A record aims at the native button inside a wrapper only when the wrapper is no tab, radio, or option, which would be pressed twice. A lone glyph is a picture only in another font family than its parent's, by each stack's own family, so "XL" in a heavier weight stays. A class ending in "-selected" or "-checked" reads as selected: found live on Myntra, whose picked size carries no ARIA state.
  13. Walls. A phone check asks to "verify the phone number", not to "sign in"; "Please check the reCAPTCHA box" and "verify that you are not a robot" are challenges; the login-wall phrases are tested.
  14. CI. The Rust job failed on clippy::assert_is_empty, new in Rust 1.99, on three test asserts; fixed, and every check below was run with 1.99.

Not fixed here, for a decision or a follow-up:

  • Decision for you: "Request ", "Book ", and "Order now" are not classed irreversible by consequence; only a planner's stop_before gates them, as it did in every live run. Gating them by wording would also ask approval for "Request a callback" and travel sites' "Book" that opens a form.
  • When a step fails, login-wall phrases are matched anywhere on the page before any rescue: a fixable failure on a page that also says "Log in to see your saved addresses" pauses for a person instead.
  • Minor: row names read textContent (screen-reader-only text can enter one); a card's ordinal restarts with each reading; a date control without a year matches any year; the slot words "to" and "address" give plain boxes a place box's two late looks; the 10 s reading cap also bounds the tree snapshot; the Jev request limit is not per provider.

Related issue

None in this repo. This is part of tinyhumansai/openhuman#7000, getting the browser agent working end to end.

API or behavior changes

No wire changes, and no contract version bump.

Engine

  • enter may press one suggestion row after typing: Jev's pick at 0.5 or more, or a lone row reading exactly as typed.
  • MAX_REQUEST_BYTES is 48,000, and a larger request is asked in parts.
  • The decision journal event gains parts, and its request_bytes is the largest part.
  • wait_for fails early on a found-nothing page.
  • Jev retries default to four, at 1, 2, 4, and 8 s. This updates the doc of JevConfig.max_retries.
  • The task controller reports reads into declared variables.

Browser

  • Sight no longer reads a region with only a tab stop, or a button holding a text box, as one control.

Core

  • human_needed has more login-wall wording, and the bare word "reCAPTCHA" is no longer a gate.
  • consequence has six more phrases.

Bus

  • FLOW_GUIDE has two clarified rules.

Round 2

  • Browser: every click goes through press_element (same tab, location permission, covered retry, link follow); page readings time out at 10 s; a timed-out navigation to a drawn page succeeds.
  • Engine: FlowRun keeps what is in front in front.rs; do steps cap repeated presses and copies; enter fails when it typed nothing; plain typing steps run as enter; new thresholds MAX_REPEAT_PRESSES, LAYER_COVERS, FRONT_CONTROLS, STEADY_HOLD/STEADY_CHECKS, LATE_LOOKS.
  • Core: result_families adds runs of single-control cards and drops bare-marker lists (BARE_CHARS, CARD_LINK_CHARS); two more login-wall phrase families.
  • No wire changes and no contract version bump.

Round 3

  • Browser: an unlabelled table row's container label carries its cells' words (row "01 Available" #1, still ending in its ordinal); a merged wrapper's record id and box are its inner native button's; aria-hidden text that takes no pointer events is read.
  • Engine: when nothing serves a do step and the task's dialog is in front, one more grounding question asks what answers the dialog; a covered control inside a dialog is offered while a dialog is in front; read offers the checked or selected items of one list as one source; a rescue's covers counts only what its guidance does before rerunning the failed step (refined in the review); new thresholds LIST_LEAN and LIST_LEAD.
  • Core: result_families drops a list split from another's cards.
  • Bus: FLOW_GUIDE gains the choice-condition and booking-button rules.
  • No wire changes and no contract version bump.

Review

  • Engine: a plain typing step is read before substitution; only variables a step writes are reported as reads; collected reads outrank a flow's declarations; the task's dialog is handed back after a step works in it; the dialog fallback is browser-only, dialog-only, and never commits; copies are struck per list and per press, after a change; enter confirms a slot's opener with Jev (new OPENER_FLOOR) and re-asks for a place box's row instead of pressing the closest; a login-only stop_before is done at once; rescue covers counts what the guidance does; requests cut an oversized screen before splitting and charge the budget per part.
  • Browser: presses retarget only the pressed link or form and same-site window.open; location is granted only in a module-launched throwaway browser; link following skips expands, downloads, in-page controls, pages already leaving, and shown dialogs; sight's pointer-events, React-handler, aiming, glyph, and selected-class rules as in 58.
  • Core: two phone-check phrases name their own action; two robot-challenge phrases.
  • Bus: the guide puts a choice's kind in a pick's from, and never plans a login stop_before.
  • No wire changes and no contract version bump.

Examples

  • task_live has a new default hosted model.
  • TASK_INTERACTIVE is new.
  • follow takes an optional Person.
  • Reads are printed, and records.json is always written.

Validation

Commands actually run, on macOS (arm64):

  • cargo fmt --all -- --check: pass

  • cargo clippy --all-targets --all-features -- -D warnings: pass

    • Run as --workspace --exclude tinycomputer-accessibility. That crate's macOS-only code has clippy failures that predate this PR and that Linux CI never compiles.
  • cargo build --all-targets --all-features: not run as its own step. The clippy and test builds compiled every target.

  • cargo test --all-features --workspace --exclude tinycomputer-accessibility: 1057 passed, 0 failed (round 1)

  • Round 2, at 428fb3be:

    • cargo fmt --all -- --check: pass
    • cargo clippy --workspace --exclude tinycomputer-accessibility --all-targets --all-features -- -D warnings: pass
    • cargo test --all-features --workspace --exclude tinycomputer-accessibility --exclude tinycomputer-examples: 1022 passed, 0 failed (917 unit and integration tests, 105 doc tests)
    • the tinycomputer-examples tests were not re-run in round 2: the disk filled while linking its binaries. That crate's code did not change in round 2.
  • Round 3, at b054d7d5:

    • cargo fmt --all: applied, then clean
    • cargo clippy --all-targets --all-features --workspace --exclude tinycomputer-accessibility -- -D warnings: pass
    • cargo test --all-features -p tinycomputer-core -p tinycomputer-browser -p tinycomputer-engine -p tinycomputer-bus -p tinycomputer: 855 passed, 0 failed
    • the tinycomputer-examples and tinycomputer-accessibility tests were not re-run: the disk is nearly full, and neither crate changed in round 3.
    • The round-3 sight changes were checked in Chromium against local copies of BookMyShow's seat-table markup (aria-hidden cells with pointer-events: none, a gridcell wrapping a left-aligned button): rows read row "01 Handicapped" #1, each "Select" record points at its BUTTON, and an aria-hidden paragraph under an overlay is still dropped.
  • Review, at 35e46635, with Rust 1.99 (CI's stable):

    • cargo fmt --all -- --check: clean
    • cargo clippy --all-targets --all-features --workspace --exclude tinycomputer-accessibility -- -D warnings: pass
    • cargo test --all-features --workspace --exclude tinycomputer-accessibility: 1090 passed, 0 failed, the examples crate included (the accessibility crate's paste test types real keystrokes on macOS)
    • GitHub CI on 35e46635: Minimum supported Rust version, Docs, and Supply chain pass; Rust was still running when this was written (its run on b054d7d5 failed on the 1.99 lint, fixed here)
    • The review's sight changes were checked in Chromium against local pages: the seat table's markup, a React-wired card with its own "Add" and an onMouseDown-only carousel track, the earlier icon-glyph pages plus "XL" and "M" in a heavier weight of the same font and "OK" with a symbols font in its stack, a store's size buttons marked by class, and a full-size pointer-events: none iframe (no longer read as a frame in front).

The round-2 sight changes were each checked in Chromium against local pages built like the live cases (naming, steppers, rows of product cards, carousel dots, icon glyphs, sprites, React-wired buttons, new-tab links).

The sight change was also checked in a Chromium page against a local copy of the popover the live run hit:

  • the old sight.js returned one button "Select a location for delivery … 560001, Bengaluru, Karnataka MG Road…";
  • the new one returns the three rows as separate buttons, while a pointer-styled carousel slide stays a button.

The gated live_* sight test was not run: this machine has no Chrome where agent-browser looks.

Live runs

Headed Chrome on the Tiny Humans route, with the module built at each stage:

Site / run Result Fixed here
Uber fare estimate both places picked from their suggestions, then needs_human at "Log in to see ride options" (~40 s) 1, 8
Myntra, cheapest white sneakers, size 9 reached PLACE ORDER (payment checkpoint), with 2 rescues the rescues: 11, 12; the missing reads: 7, 15
Amazon, BookMyShow ended by HTTP 502 on oversized knockouts 4, 5
Amazon, re-run 2 (after 1's floor and 11's filter rule) reached "Proceed to Buy", but with another brand's earbuds (pick by "lowest price"), and only after 4 rescues: "Go" was lost with the second part's answers 4 (ballots), 11 (pick wording)
Amazon, re-run after 4 and 5 no 502s: 718 calls, 0 failed, 11 decisions split (largest part 47.6 KB). It failed later: a suggestion pick changed the search to "…141 anc", and a "skip Sponsored items" step had nothing to do 1 (floor), 11 (filter rule)
BigBasket wrong delivery area (popover pressed), then "No Results Found" ×3 2, 3, 6
Ola suggestion rows inside shadow DOM, not readable not fixed here (below)
BlazeDemo / Selenium (model change) 84 → 44 s, 65 → 29 s 13

Round 2 (module rebuilt between fixes):

Site / task Result
Zepto: search Maggi, add 2, cart total ✅ cart with 2 packets, ₹40/₹50 (twice, 93–102 s)
Flipkart: first boAt Airdopes 141 rated 4★+, title/price/rating, cart ✅ payment checkpoint ("Place order"), 2 rescues
Lenskart: first blue-light glasses, cart ✅ payment checkpoint ("Proceed To Checkout"), ₹500, 2 rescues
Blinkit: current location, Amul Taaza 1 L ×2, cart ✅ end to end (location detected, quantity 2, ₹95 grand total, "Login to Proceed"); one rescue took the 200 ml pack, since fixed in the rescue protocol (35)
Uber: book a cab, cheapest car, stop before Request ✅ after the person signed in: 9 options extracted, Uber Go AC ₹329.64 picked, gated at "Request Uber Go AC"
BigBasket: current location, Amul Taaza 1 L ×2 ⚠️ right product added once; the basket needs an OTP login the task forbids (now a needs_human pause, 34); its place popover has no current-location option
BookMyShow: show on Wed 7 Oct, 2 seats ❌ reaches the show timings with the format answered; the date strip is not re-chosen after the format dialog, and a pick opens a cinema card instead of a show time
Ola: pickup and drop ❌ the drop was typed over the pickup and the late suggestion rows were not picked; fixed in 31–32, not yet re-run

Round 3 (module rebuilt between fixes; every approval prompt was answered by the person running the test):

Site / task Result
Uber: cheapest car, stop before Request ✅ after sign-in: all 9 options with fares, Uber Go AC picked, gated at "Request Uber Go AC"; the pick had pressed the already-selected option, opening a fare breakdown, fixed (36)
BigBasket: Amul Taaza 1 L ×2 ⚠️ the right milk added and raised to 2 cleanly; the basket needs an OTP login, so it pauses as needs_human
Ola: MG Road Metro → airport, cheapest car ✅ pickup and drop set, ride types listed, gated at "Confirm & Book" (declined). Fares need a login ("Please log in to check exact prices"); it chose Auto, not a car, and reported the cards' lines (42, 46, not yet re-run)
BookMyShow R11 ❌ reached the accessible seat table; the second seat's "Select" was struck off (44); a Handicapped seat was pressed (40)
BookMyShow R12 ❌ further: show time kept (41), "Select Seats" pressed in-step; six seat presses missed their buttons (40)
BookMyShow R13 ❌ date, format, show time and seat count in-step, no rescue; seats A-02 and A-03 chosen, cheapest section ₹200, "Pay ₹400"; reading the seat numbers failed (aria-hidden cells, 40)
BookMyShow R14 ❌ seats A-03 and A-04 chosen; seat numbers now read, and the combined source showed "03 Available; 04 Available", but Jev took a single number (0.26); the source now says plainly that these are the checked items (43), not yet re-run

Review pass, on the build with the sign-in rule (53), before the other review fixes:

Site / task Result
Amazon: first boAt Airdopes 141 rated 4★+ ✅ "Proceed to Buy (1 item)", no rescues. Every Airdopes 141 result now shows 3.8★, so the first non-sponsored result rated 4★+ was "boAt Nirvana Ion", as the task's words ask
Myntra: white sneakers, size 9 ✅ PLACE ORDER, total ₹304, one rescue at the size step: the chosen size showed no state, fixed in 58
Flipkart, Zepto, Lenskart, Blinkit, BigBasket, Uber, Ola, BookMyShow being re-run one by one on 35e46635; results follow in a comment

Known and not fixed here:

  • Ola: its rows sit inside shadow roots. A fix to agent-browser's snapshot (marking pointer and click elements inside open shadow roots, resolving them with a piercing DOM.getDocument) is in the local vendor checkout but not in this PR: it needs its own commit on the tinyhumansai/agent-browser fork, then a gitlink bump here.
  • BookMyShow: the seat read with the clearer combined source is not re-run yet. A planned pick of "the shows" still opens a cinema card; the seat step then presses the show time itself.
  • BigBasket: the header's basket icon is not read as a control.
  • Warning noise: the "failed to prepare iframe session" warnings belong in the agent-browser fork.

Tests

  • flow_tests/suggestion_tests.rs, with a simulator ride form (flow_tests/places.rs):
    • picks a suggestion, with Jev and directly;
    • a plain field asks nothing;
    • a private text is never offered;
    • a panel that strings its rows together is never pressed.
    • Without their fixes, the first and last fail as live did: the box emptied, and the wrong place was picked.
  • flow_tests/split_tests.rs:
    • parts stay under the cap, each with the whole state;
    • every question is asked once, in order;
    • a 400-row knockout goes out in parts and still clicks its target. This test fails with splitting disabled.
  • step_kinds_tests::a_wait_for_stops_when_the_page_says_it_found_nothing
  • resolve_tests::jev_calls_ride_out_a_provider_outage_unless_told_otherwise
  • output_tests::a_value_read_into_a_declared_variable_is_reported: fails without the fix, with "no entry found for key".
  • safety_tests:
    • login walls, the reCAPTCHA badge, and the challenges;
    • the new payment and irreversible phrases.
  • sight_tests/live_tests::live_panels_and_regions_leave_their_rows_to_be_read: gated on TINYCOMPUTER_LIVE_BROWSER=1.
  • Examples task_tests.rs:
    • what a person sends a paused task;
    • read lines;
    • prompt sentences.

Round 2 adds:

  • step_kinds_tests::a_plain_step_that_types_is_read_as_the_enter_it_means
  • pick_tests::a_first_with_a_condition_walks_the_list_in_order
  • choose_tests::a_day_in_a_strip_of_dates_is_found_by_its_short_label
  • suggestion_tests::a_search_box_takes_only_the_same_search_and_a_place_box_its_reworded_rows
  • safety_tests::walls_only_a_person_can_pass_are_named: OTP sign-in and "Sign up or Log in" walls; a header's "Login/ Sign Up" is none.
  • Browser surface tests now allow the same-tab script and the pointer move before a covered retry (Fake::evaluated_besides_keeping_the_tab, Fake::pressed_by_position).
  • Not yet covered by simulator tests: the late suggestion looks and the steady wait_for. They were exercised in the live runs above. The dialog tracking in front.rs gained a test in the review.

Round 3 adds:

  • pick_tests::a_pick_the_page_already_has_selected_is_not_pressed_again
  • suggestion_tests::a_place_box_takes_its_closest_row_when_no_suggestion_clearly_fits (replaced in the review, below)
  • do_loop_tests: a step choosing several items presses each one's own copy, and one adding a count of one item does not (the copy rule); several items are asked for by a choosing verb and a counted plural; a step finding nothing to press answers the task's dialog; a control the dialog's own bar covers is pressed and one behind it is not. Each simulator test fails with its fix switched off.
  • rescue_tests::guidance_that_ends_by_running_the_failed_step_again_covers_nothing_after_it
  • groups_tests::a_list_of_another_lists_cards_split_line_by_line_is_dropped
  • pick_tests::a_list_not_clearly_chosen_is_the_one_jev_leaned_to_when_it_leads_clearly
  • step_kinds_tests::the_items_chosen_together_in_one_list_read_as_one_source

The review adds, each run with its fix switched off and seen to fail:

  • step_kinds_tests: a_plain_typing_step_types_a_value_that_names_a_fact_as_written; new cases in a_plain_step_that_types_is_read_as_the_enter_it_means; a_stop_before_names_only_signing_in_or_something_irreversible; a_stop_before_signing_in_lets_the_flow_go_on; a_value_read_before_a_rescue_outranks_the_flows_empty_declaration
  • output_tests::a_variable_defined_from_a_fact_is_never_reported_as_a_read
  • do_loop_tests: a_dialog_the_task_worked_in_is_in_the_way_of_the_next_step; answering_the_task_dialog_never_presses_what_commits; a_pressed_controls_copies_are_only_on_the_other_items_of_its_list
  • enter_tests::a_control_named_by_the_slot_opens_its_box_only_when_jev_agrees
  • suggestion_tests::a_place_box_asks_again_for_the_row_naming_its_place_in_other_words
  • pick_tests: "First AC", "first class", "first to depart", "first alphabetically" are no conditions
  • rescue_tests::guidance_that_does_the_later_steps_before_rerunning_the_failed_one_covers_them
  • split_tests::a_screen_that_fills_most_of_a_part_is_cut_rather_than_asked_once_per_question
  • records_tests::first_and_last_alone_are_the_lists_own_order; safety_tests: the new wall and challenge phrases
  • browser operations_tests: only_a_control_asking_for_the_place_is_read_as_one; two_addresses_of_one_page_are_one_place; location_is_granted_only_in_a_browser_the_module_launched

Documentation

  • Engine flow docs:
    • docs/crates/tinycomputer-engine/flow/filling-forms.md
    • docs/crates/tinycomputer-engine/flow/step-kinds.md
    • docs/crates/tinycomputer-engine/flow/voting-and-briefing.md
    • docs/crates/tinycomputer-engine/flow/budgets-and-switches.md
  • Jev runtime and harness:
    • docs/crates/tinycomputer-engine/jev-runtime.md
    • docs/technical/jev-harness.md
    • docs/technical/jev-journal.md
    • docs/technical/jev-questions.md
    • docs/technical/decision-thresholds.md, which also gains the missing MOST_SUGGESTIONS and OPTION_EXTRA_WORDS rows
  • Tasks, sight, and safety:
    • docs/technical/tasks.md
    • docs/technical/specs/browser-sight.md
    • docs/safety-and-privacy.md
  • Configuration and examples:
    • docs/crates/tinycomputer/configuration.md
    • docs/crates/tinycomputer-examples/live-tasks.md
    • .env.example
  • Planner guide: crates/tinycomputer-bus/src/flow/guide.md
  • Round 2: docs/technical/decision-thresholds.md (seven new constants), docs/crates/tinycomputer-engine/flow/step-kinds.md, docs/crates/tinycomputer-engine/flow/filling-forms.md
  • Round 3: docs/crates/tinycomputer-browser/sight.md, docs/technical/specs/browser-sight.md, docs/crates/tinycomputer-engine/flow/step-kinds.md, docs/crates/tinycomputer-engine/output.md, docs/technical/specs/task-rescue.md, docs/technical/decision-thresholds.md (LIST_LEAN, LIST_LEAD, and the copy rule's exception)
  • Review: docs/crates/tinycomputer-engine/flow/the-do-loop.md (the task's own dialog, presses and copies), filling-forms.md (search versus place boxes, the place re-ask, the opener, typing steps), step-kinds.md (the login-only stop_before, "first" with a condition), docs/technical/decision-thresholds.md (OPENER_FLOOR, the missing MAX_IDLE_SCROLLS, the copy and layer rows), jev-harness.md, jev-questions.md, specs/task-rescue.md, docs/crates/tinycomputer-browser/sight.md and specs/browser-sight.md (naming order, chosen and hidden states, React controls, what a press adds to a page), docs/safety-and-privacy.md (the location grant, new tabs), docs/rescue.md, docs/crates/tinycomputer-core/prices-times-and-dates.md

Checklist

  • The change is focused on one logical change: it gathers the fixes from one round of live testing, each in its own commit and each reviewable on its own.
  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

With TINYHUMANS_TOKEN and no model named, task_live planned, rescued and
shaped on the gateway's agentic-v1 tier. That tier reasons for 25 to 45
seconds over a small rescue, and on a real one it ran past the module's
120-second rescue limit; planning alone took about 40 to 50 seconds of a
BlazeDemo or Selenium run. The default is now
openrouter/deepseek/deepseek-v4-flash, the managed default OpenHuman runs
its own hosted work on, which answers the same rescue in 7 to 13 seconds.
A model named through TINYCOMPUTER_PLANNER_MODEL, _RESCUE_MODEL or
_OUTPUT_MODEL still wins.
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 08804e41-1814-495d-8417-a759366b181d
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@YellowSnnowmann
YellowSnnowmann marked this pull request as ready for review October 6, 2026 06:23
@tinysweeper

tinysweeper Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 657fffb02f39. the review of #76 did not finish within 900s

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0009 · 72,330 in / 4,357 out · 6,212 cached (9%)  · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0004 · 36,336 in / 2,099 out · 6,148 cached (17%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0003 · 21,020 in / 1,112 out · 0 cached (0%)      · gpt-5.6-luna
tests:       $0.0001 · 5,458 in  / 231 out   · 64 cached (1%)     · glm-5.3-flash
description: $0.0001 · 5,585 in  / 94 out    · 0 cached (0%)      · glm-5.3-flash

Comment thread crates/tinycomputer-examples/src/bin/task_live/main.rs
A location, city, or airport box lists matches under itself as you type
and keeps the text only once one of them is chosen; moving the focus on,
or pressing Escape on the list, drops it. enter typed the text, saw it
read back, and moved on, so the next step took the open list for
something in the way, pressed Escape, and the box emptied (live: a ride
site's pickup and dropoff boxes both ended blank and the task failed
after five rescues).

Once a text has arrived, enter now looks again and, when rows appeared
that were not on screen before it was typed, picks the one matching it:
rows that mention the text first, a single match pressed without asking,
otherwise Jev picks among at most 12 or answers that none fits. Without a
mention, only new rows drawn as a list's rows are offered, so a button
that appeared beside the box is never taken for a suggestion. A private
text is never offered, and a field that opens no list asks nothing.

The simulator gains a ride form whose boxes behave this way; the new
test fails without the change with the pickup box emptied.
A step that failed in front of a login wall becomes needs_human only when
the page shows one of the gate phrases, and those knew only "sign in to
continue" and its kin. A ride site's dialog read "Log in to see ride
options" and "please take a moment to quickly log in or sign up", so the
task spent its rescues and failed instead of handing over to the person.

The gate phrases now also cover a wall that names what it hides ("log in
to see", "sign in to view"), the dialog's call to action ("log in or sign
up"), "please log in", "you must be logged in", and "login required". A
header's bare "Log in" and "Sign up" links, with no "or" between them,
still read as no wall.
An invisible reCAPTCHA puts a frame titled "reCAPTCHA" (and a "protected
by reCAPTCHA" notice) on every form it guards, and the gate list matched
the bare word as a captcha to solve. Any step that failed on such a page
became needs_human with "solve the captcha" when there was none: live, a
cab site's pickup page was handed to a person with no captcha in sight.

reCAPTCHA now counts only by its challenge: "complete the reCAPTCHA",
the challenge frame ("reCAPTCHA challenge expires in two minutes"), and
the image grid ("select all images", "select all squares"); "I'm not a
robot" already did, and hCaptcha's "I am human" joins it. A bare
"captcha" still counts, so a text captcha keeps pausing for a person.
task_live ended the run at every pause only a person can answer: an
approval, a login or captcha, a detail it was not given, and a payment
page. It then closed the browser the task had opened, so none of those
steps could be tried standalone.

With TASK_INTERACTIVE=1 the run waits for the person at the terminal:
approve or decline an irreversible action, get past a login or captcha in
the browser window and press Enter, type a detail the answers lack, and
finish on a payment page before the browser closes. End of input answers
no, so a run with nobody at the terminal stops rather than waits.

follow takes the person as an option; task_fixture passes none and is
unchanged. What the person sends a paused task is a pure function
(reply), tested with a scripted person.
@YellowSnnowmann YellowSnnowmann changed the title fix(examples): run task_live's hosted roles on the managed default model fix: autocomplete picks, login walls, faster model, interactive runner Oct 6, 2026

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0093 · 742,662 in / 52,997 out · 112,689 cached (15%) · gpt-5.6-luna, glm-5.3-flash, deepseek-v4.1-flash
critique:    $0.0056 · 430,663 in / 31,300 out · 67,654 cached (16%)  · gpt-5.6-luna, glm-5.3-flash
security:    $0.0030 · 224,473 in / 14,886 out · 23,275 cached (10%)  · gpt-5.6-luna
tests:       $0.0004 · 45,230 in  / 3,936 out  · 20,480 cached (45%)  · glm-5.3-flash, deepseek-v4.1-flash
description: $0.0001 · 20,300 in  / 795 out    · 1,280 cached (6%)    · deepseek-v4.1-flash

Comment thread .env.example
Comment thread .env.example
Comment thread crates/tinycomputer-examples/src/bin/task_live/main.rs
Comment thread crates/tinycomputer-examples/src/bin/task_live/main.rs
Comment thread docs/crates/tinycomputer-engine/flow/filling-forms.md
Comment thread crates/tinycomputer-engine/src/agentic/flow/flow_tests/simulator.rs
Comment thread crates/tinycomputer-examples/src/bin/task_live/main.rs
Comment thread crates/tinycomputer-examples/src/task/person.rs
Comment thread crates/tinycomputer-engine/src/agentic/flow/steps/suggestion.rs
Comment thread crates/tinycomputer-examples/src/task/person.rs
@YellowSnnowmann
YellowSnnowmann marked this pull request as draft October 6, 2026 08:33
A store's delivery-area popover was read as one button whose name held
every suggestion row, so after a pincode was typed it was the only new
element mentioning the text and was pressed without asking. The press
landed on the row at the panel's middle and set another area than the
one typed. A label that lists more than the typed text is no longer
offered as a suggestion; the box keeps the text as typed.

Also lists MOST_SUGGESTIONS and OPTION_EXTRA_WORDS in the decision
thresholds table.
Sight made any element with a tab stop, a pointer cursor, or a click
handler a button, and then skipped every element inside it. A store's
delivery-area popover took a tab stop, so it was read as one button
named with its heading and every suggestion row, and the rows could not
be pressed. A tab panel's tab stop hid its rows the same way.

A region a page marks as holding controls (menu, list, grid, tab panel,
dialog, a landmark) is no longer a control itself, and a button or link
that holds a box to type in is read as a panel. The rows inside either
are read as controls of their own.
A grounding knockout over a long results page carried 16 to 19 group
questions, 73 to 79 KB of JSON, and the Tiny Humans gateway refused it
with HTTP 502 on every attempt, which ended the task (Amazon, BookMyShow).
Fitting could not help: it shrinks only the state, which was under 9 KB,
and it dropped every question's brief on the way.

A request over the cap is now cut by its questions into parts, each with
the whole state, asked at once; Jev evaluates each question on its own,
so the answers merge back by question id. Each part is then fitted as
before. The cap drops from 100 KB to 48 KB: across 7,068 live calls the
largest that passed was 57 KB (23,600 tokens) and the smallest refused
was 68 KB. The decision journal event gains `parts`, and its
`request_bytes` is the largest part.
The Jev client retries a 5xx or 429 twice, 100 ms and then 200 ms
apart. Live, a gateway's 502s ended a task after three attempts in
2.5 s. Unless the configuration sets `max_retries`, a Jev call now
retries four times, waiting 1, 2, 4, then 8 seconds (or what the
provider asks for), about 15 s in all. Building the client
configuration moves into `client_config` so the policy can be tested.
"Place your order", a store's last button past the payment page, did
not match the whole-word phrase "place order", so it read as
reversible; neither did "Confirm order" or "Complete order". They and
"Proceed to pay" now count as payment. "Confirm ride" and "Confirm
pickup", which send a driver, now need approval. "Book" alone stays
reversible: on travel sites it opens the traveller form.
A store's search landed on a page titled "No Results Found", and
`wait_for` checked it ten times, about 50 s, three times in one task.
Each rescue was told only "still not true after 10 checks" and guessed
at the query's wording. A page whose title or visible text says it
found nothing ("no results", "0 results", "no products found", ...) on
two checks in a row now fails the step at once, naming the phrase, so
the rescue sees the real reason. The note names a phrase from our own
list, never the page's text.
The planner declares each read's variable up front, empty ("total":
""), and the task dropped every variable the flow declared as the
caller's own input. So the brand, price, and bag total a task was told
to report never reached its records, its answer, or TaskReport; only an
undeclared pick came back. A variable now counts as read when the run
changed it from what the flow gave it.
task_live wrote records.json only for a finished task and printed
nothing of what was read, so a run stopped at the payment checkpoint
showed its steps but not the values it was asked to report. It now
prints one `read <name>: <value>` line per variable and writes
records.json from the report however the task stopped.
The guide said choosing by a criterion is what `pick` is for, so a
planner wrote `pick` from "the available size options" by "closest to
UK 9". A `pick` ranks result cards to open, and over option buttons it
failed, costing a rescue that then chose "9". The guide now says an
option group (size, colour, quantity) is a `choose` of the shortest
label the page is likely to show.
A planner wrote `wait_for` "the product grid has updated after
sorting". The flow judges a condition on the current screen, and at
the deep level once more over the screen alone, without the history:
there, a change cannot be seen (0.64), which pulled each check under
the bar (0.70 against 0.75) on a list already sorted. Ten checks and a
rescue followed. The guide now says to name what shows once the change
has happened.
The payment prompt ran two sentences together ("...paying is left to
you The page stays open for you"), and a login prompt would have shown
two full stops for a reason that already ended in one.
@YellowSnnowmann YellowSnnowmann changed the title fix: autocomplete picks, login walls, faster model, interactive runner fix: browser agent fixes from live tests on six sites Oct 6, 2026
The previous commit stopped reading any element with a region role
(menu, list, group, tab panel, dialog) as a control. Some pages make
such an element the click target itself, such as a carousel's slide
(`role="group"` with a pointer cursor), and it would have lost its only
control. A region is now left out only when its tab stop is all that
made it look pressable; with a pointer cursor or a click handler it is
still a button.
Live, after "boAt Airdopes 141" was typed into a store's search box,
its completions ("... 141 anc", "... 141 elite", ...) were offered as
suggestions, and Jev's pick of "... 141 anc" at 0.43, a near tie with
"none fits" at 0.42, was pressed: the search itself changed, and the
task failed further on. The right places on a ride site came at 0.58
and 0.78.

A suggestion Jev picks is now pressed only at 0.5 or more
(SUGGESTION_FLOOR), and a lone matching row is pressed without asking
only when it reads exactly as typed; one that says more is Jev's to
pick.
A planner turned "skip Sponsored items" into a `do` step of its own.
There is nothing on screen to do for a filter, so the step spent eight
turns and a rescue. The guide now says a filter on which result to
take belongs in the `pick`'s `from` or `by`.
`vote::ballots` read its question ids from the first answered framing's
request. Once a request is split by its questions, the framings of
different parts ask different questions, so every part but the first
lost its answers. Live, a search's "Go" button won its knockout group
at 0.98 in the second part, the group's answer was dropped, and the
knockout had no winner: the step pressed nothing for three turns and
failed, and so did two rescues after it.

The ballots now cover every question any framing asked. The split test
asks for a target in the last group, which goes out in a later part;
it fails without this fix, as live.
Asked for the first boAt Airdopes 141 rated 4 stars or more, a planner
wrote a `pick` from "the search results that are not sponsored and have
a rating of 4 stars or more" by "lowest price". A store lists other
brands beside the one searched for, so the pick took the cheapest of
them, a Fire-Boltt pair, and added it to the cart. The guide now says a
pick's `from` names the item asked for and its `by` is the task's own
criterion, "first" when the task says first.
A store's basket opened a "Login/ Sign up Using OTP" dialog, and a ride
app's "See prices" led to "Sign up or Log in". Neither paused for the
person, so rescues looped on a dialog only a person can pass. Both are
login walls now, with "enter OTP" and "verify your mobile/phone number".
A header's bare "Login/ Sign Up" still says no more and is no wall.
Sight is the page reader every browser step sees through. Live, on
Indian store and ride sites, it misnamed or missed what a person sees:

- a box is named by its own label, aria label, placeholder, or title
  before the words beside it; nearby words only name a box that says
  nothing itself, and a divider's "OR" never does. A search box read as
  "Location not set" and a location box as "OR", and no enter found them;
- an unmarked picture button beside a bare count is "decrease" (before
  it) or "increase" (after it); a store's quantity buttons had no name;
- plain boxes repeated as siblings of one tag and class, each with a
  link or button, a word, and links to at most two places, are cards,
  numbered across rows so a grid of rows of four reads as one list; a row
  holding counted cards is their row, not a card;
- a lone glyph drawn in an icon font, or in a font other than its
  parent's (digits, currency, and + - x < > excepted), is a picture, not a
  name: a store's search and close buttons read as "p" and "!";
- an icon's sprite symbol and picture file name ("#icon-cart",
  "cart.svg") name it too;
- a small element whose React props hold a click handler is a button: a
  store's "Add to cart" and "Buy now" were plain words, with nothing to
  press.
Every click on the browser surface now goes through press_element:

- links and forms aimed at `_blank` are pointed at the page's own tab
  first, and for two seconds a script's window.open(url) opens there:
  product links opened new tabs the session never read;
- a control that asks for the person's location ("use my current
  location", "detect my location") gets the session the geolocation
  permission first, standing in for the browser's Allow bubble the agent
  cannot press;
- a click refused as covered is tried once more with the element in the
  middle of the window and the pointer moved to the corner: a product
  photo's hover zoom covered "Add to cart" twelve times, and sticky bars
  covered rows scrolled under them;
- a content link (four words or more) whose press left the page where it
  was after 400 ms, with no dialog opened, is followed to its address: a
  card's carousel took six presses of its product link.

A navigation that timed out counts as open when the page shows the same
host and path with words drawn, and one reading of the page, by sight or
as a tree, is capped at 10 s instead of 30 s each. Pauses on the surface
are plain sleeps, since its calls block by contract.
FLOW_GUIDE gains the patterns live runs on stores, ride apps, and a
cinema showed the planner and its rescuer kept missing, all generic:

- a dialog's headings group its buttons and are not answers;
- a choice made by picking a suggestion is set: no step to confirm it;
- a delivery store: set the place first when the task names one; with
  "use the current location if asked", only where the page asks and
  offers that control, and never wait for a place to be set;
- to buy more than one, add the item first and raise its count next;
- on a page that lists results as the query is typed, the search step
  finds its work done;
- seats are a plain step naming the section and count, never a pick;
- a day in a strip of dates is a choose once the strip shows;
- a stop_before comes where the flow would take that action, and never
  for a login the task did not ask for.
The do loop, from live runs on stores, ride apps, and a cinema:

- A dialog the run's own press opened is its next stage. front.rs keeps
  what is in front across steps and rescue runs: a sheet, or a layer that
  covers three more controls than before the press, on the same page, and
  not opened by typing (a search box's suggestions are the field's). Such
  a dialog is never cleared by attention, an obstacle dismissal, Escape,
  or an undo; while it is in front, controls it covers are not offered as
  moves, nor its close control unless the step says to close. A browser
  run's first look takes a dialog in front as the task's (a rescue's
  inheritance); an application's own opening alert is still cleared.
- A step that fails in front of that dialog names its own controls, so a
  rescue answers with one of them: rescues kept choosing a language
  heading above a format dialog's buttons.
- A control, or a key, pressed three times in a step is not pressed
  again (a toggle's open and closed looks count as one control). Once a
  named control is pressed, its copies on other items are struck off for
  the step unless it says all, every, each, or both: a step adding two
  packets of one milk pressed "Add" on six products. An undo lifts them.
- Attention leaves a layer open when a control in front fits the step as
  well as anything behind it: a location dialog was escaped by the step
  that meant to press its "Use my current location".
- "allow selection" and "save my choices" are consent-card closers.
- The move option says a control scrolled out of view can be pressed;
  the judge called a step stuck rather than press one.
- A step that stalls says its work may already be done, so a rescue
  skips a search button a live search does not have.
…osely

Steps, from the same live runs:

- A plain step that asks for typing ("enter 560001 into the pincode
  field", "fill in the pincode with 560001") runs as the enter it means:
  a do step cannot type. An enter that typed nothing now fails instead
  of reporting done; it first presses a control named by the slot's own
  word (a search link for the "search box"); and a box one slot was
  typed into is never another's, by ref or by the text it holds: a ride
  app's drop was typed over its pickup.
- A search box takes only a suggestion that is the same search ("Show
  all results for ..."), pressed at once. A place box looks twice more
  for late suggestions, and there a row already showing counts when it
  matches, as does a pressable box sharing half the typed words.
- A pick by "first <condition>" walks the list in page order for the
  first item that meets the condition: judged at once, it took the fifth.
- wait_for holds on a settled screen judged at 0.65 or more three checks
  in a row; more "found nothing" phrases end it early.
- A date option matches a strip's "WED 07 OCT": leading zeros, short
  months, and no year unless shown.
- A read takes the text holding the value itself, the shortest.
- The rescuer is told a delivering store finds nothing until its place is
  set, to keep the task's size and variant words, and never to stop
  before a login the task did not ask for.
decision-thresholds.md gains MAX_REPEAT_PRESSES, LAYER_COVERS,
FRONT_CONTROLS, STEADY_HOLD/STEADY_CHECKS, LATE_LOOKS, BARE_CHARS, and
CARD_LINK_CHARS. step-kinds.md covers a pick by "first <condition>", the
card opener order, and wait_for on a settled screen; filling-forms.md
covers search and place boxes, an enter that typed nothing, the control
named by a slot, one box per slot, and plain steps that type.
Live, a ride app's cheapest car was selected by default. The pick
pressed it again, which opened a fare breakdown over the "Request"
button, and the step that stops before requesting could not find it.

A pick whose item's control shows selected or checked now takes the
item without pressing it: its text goes into the variable and the step
is done, noting "already selected". The simulator can mark a result
card's control selected, and a test pins that nothing is pressed.
Live, a store's "Add to basket" became a − 2 + stepper and the basket badge read 1, yet the judge and the rescuer took the item as not added and pressed on. The guide now says such a stepper means the item is in the cart with that count.
…tems

Live, a store's search dropdown stayed open over its basket button after
an item was added from it; Escape left it there, and every press of the
basket was refused as covered. The retry of a covered click now blurs the
focused text box first, which closes the suggestions most boxes hold
open, before it centres the element and moves the pointer off.

The same run's last rescue pressed "Add" on another size of the milk it
had already added, reading the open row as not yet added. The rescue
protocol now says an add button that became a minus, count, plus stepper
is in the cart with that count: never add it again, nor another size.
Live, a ride app listed "MG Road Shivaji Nagar Bengaluru" and "MG Road
Shanthala Nagar …" for the pickup "MG Road Metro Station, Bengaluru".
No row named the station, Jev rightly found no clear fit, the text was
left as typed, and a pickup that is only typed is no pickup: no ride
type ever showed. A place box (not a search box) now takes the row
sharing the most of the typed words, the first on a tie, when no
suggestion clears the floor. A search box still keeps its text.

The rescue protocol also says a strip of dates opens on today, so the
day asked for counts as chosen only when the strip shows it selected,
and that a time listed inside a card is pressed by a plain step, never
picked: rescues skipped a cinema's date and opened a cinema card.
Three things a seat table's accessible view showed live:

- A row with no label of its own read as `row tinyhumansai#1`, so the status cell
  beside a seat's "Select" ("Handicapped") never reached Jev. An
  unlabelled table row is now named by what its control-free cells say:
  `row "01 Available" tinyhumansai#1`.
- Each seat's `td role="gridcell"` held a native `<button>` at its left
  edge, and the two merged into one control kept on the cell, so every
  press went to the cell's middle and selected nothing. A merged
  wrapper's record now points at the native button or link inside it.
- The seat number and status cells were `aria-hidden` and took no
  pointer events, so the hit test that decides whether hidden text is
  plainly seen passed through them and every row in view lost its
  number. Landing on what holds such an element now counts as seeing it.
Live, a rescue pressed "Book tickets", chose the date again, and said its
steps also covered the step after the failed one, so the show time that
step was to pick was never picked and the next rescue was spent on it.
Guidance whose last step runs the failed step again (the same action, or
the same `do` intent in other case or spacing) does none of the steps
after it: its `covers` is taken as 0, and the rescue protocol says so.
Live, a ride app's five option cards also read as their nine lines (a
description, then an arrival time). The lines were the longer list, so
an `extract` that Jev answered with no clear list fell back to them, and
the ride names and fares were lost.

- `result_families` drops a list of one-line items that each sit inside
  one of another list's three or more cards and outnumber them: those
  cards' lines, split apart. Lists of several-field items stay.
- A list not clearly chosen goes to the one Jev leaned to when it has
  `LIST_LEAN` (0.3) or more and `LIST_LEAD` (3) times the next list's
  probability; only otherwise to the longest. Live, the cards drew 0.41
  against 0.05 for any other list.
Live, a seat table marked two seats "Selected", and a read of "the
selected seats" could take only one piece of text, none of which named
both. Where two or more controls in one list show checked or selected,
the read also offers what their cards or rows say, together, as one
source described as those checked or selected items
(`02 Companion; 03 Available`), its text masked like any other value
when values are not shown.
Pressing a named control strikes off its copies on other items, so a
step adding two packets of one milk cannot press "Add" on six products.
Live, that also struck off the second seat's "Select" in a seat table,
and "choose 2 adjacent available seats" could never choose the second.

A step whose choosing verb (choose, select, pick, tick, check, mark)
comes with a count above one before a plural ("choose 2 adjacent
seats", "select three files") keeps the copies. A count of one item to
add ("add 2 packets of milk") leads with no choosing verb, and keeps the
rule.
Two things a booking's dialogs showed live:

- A date step found nothing to press for four turns while the format
  dialog a booking button had opened offered "2D": a format serves no
  date step's words, and a rescue was spent pressing it. When grounding
  finds nothing for the step while a dialog the task opened is in front,
  the step now asks which of the controls a press reaches answers that
  dialog the way the task wants, and presses it (never the dialog's
  close control, and not remembered as the step's own control).
- A seat table's lower rows sat under its own "Pay" bar, read as
  covered, and were never offered. What is covered stays out while a
  dialog is in front only when it lies outside a dialog: a control the
  dialog's own bar covers is pressed, and the press scrolls it clear.
- A plan for "choose the cheapest car" picked by "lowest fare" alone,
  and a rescue then chose an auto. The flow guide and the rescue
  protocol now say to keep every condition the task puts on a choice:
  "the lowest fare among the cars".
- A film's, show's, or stay's page offers its dates, times, and seats
  only once its booking button is pressed: the guide says to press it in
  a step of its own before choosing a date.
…t picks

From the PR review:

- A "verify your phone number" wall told the person to "sign in"; it now
  says "verify the phone number". "Please check the reCAPTCHA box" and
  "verify that you are not a robot" are challenges a person must pass.
- The login-wall phrases this PR added are tested, and so are the
  `First` and `Last` criteria: a bare "first" or "last" is the list's own
  order, while "first product rated 4 stars or more" and "First AC" are not.
- The criteria and rescue docs say what the code does.
From the PR review of the browser adapter:

- Keeping a press in the agent's tab now changes only the pressed link or
  form, never every link on the page; a script's `window.open` is opened
  in place only for an address on the same site (an advert another site
  opens on a click keeps its own window), and the patched `open` returns
  no window, so a page closing "its" window never closes the agent's tab.
- Location is granted only in a browser the module launched on a
  throwaway profile, never in a person's own browser or profile, and
  "auto detect" (a translator's language) no longer reads as asking where.
- A link whose press went nowhere is no longer followed when the press
  was an expand or collapse, the link downloads a file or controls the
  page, the page already began to leave (a slow link), or a dialog is
  shown in front (hidden ones in the DOM aside); a follow that fails
  leaves the press's own reply.
- A page that drew another search is no page of the search asked for:
  the query counts in comparing addresses.
- Pressing again after a refusal no longer lets go of the focused box
  when the target is a row of its own list of suggestions.
- Sight: only an element that turns pointer events off itself counts as
  seen through a hit on what holds it, never a hit on the page's root (a
  modal library turns them off for the whole body); only a press handler
  (`onClick`, `onPress`) makes a React element a control, and such a
  control hides nothing pressable inside it; a record is aimed at the
  native button inside a wrapper only when the wrapper is not a tab, radio,
  or option, which select by their own state; a lone glyph is a picture
  only in another font family than its parent's, judged by each stack's
  own family, so "XL" in a heavier weight of the same font stays; and a
  class ending in "-selected" or "-checked" (a store's picked size) reads
  as selected.
From the PR review of the flow engine, each with a test that fails
without its fix:

- Secrets. A plain step that types is read from the step as written: its
  substituted text was substituted again by `enter`, so a value read off a
  page that said `${card_number}` was typed as the caller's card number.
  A variable the flow defines from a fact (`"recipient": "${email}"`) is
  never reported as a read: only a variable a step writes is.
- Typing steps. "Enter" also means going into something ("enter Reader
  mode in Safari") and types only what is quoted, data, or into a box; a
  step that does more than type ("… and press Enter") stays a plain step;
  "type in" is a verb; a lone `${otp}` is typed without "the"; a quoted
  text ends at its quote, and the split prefers the field naming a box.
- The task's own dialog. A dialog a step worked in is handed back at the
  next step, so a calendar left open is cleared again; a scroll or the
  run's housekeeping opens no dialog of the task's, and opening an
  address forgets it. Answering such a dialog when nothing serves the
  step happens only on a browser task, among its own controls, and never
  presses what commits ("Yes", "Confirm", "Pay").
- Presses. A control's copies are struck off only on the other items of
  its list, once the next look shows the press changed the screen, and an
  undo lifts only that press's copies; a scroll counts toward no press cap.
- `enter` presses a control named by a slot's word only once Jev agrees it
  shows that slot's box (`OPENER_FLOOR`), never on a shared word or the
  control's role word alone. A place box with no fitting row asks Jev once
  more for the row naming the place in other words, instead of pressing
  the row sharing most words; a search box offers only new rows, and a
  place box whose own words say "Search for area…" takes a place.
- Picks. "first X" is a condition only with a condition word (rated,
  under, with, …): "First AC" is a name. The flow guide and rescue prompt
  put a choice's kind in the pick's `from`, where the ranking checks it.
- A `stop_before` that names only signing in gates nothing (a login wall
  pauses by itself); signing up or creating an account stays gated.
- Rescues cover only the later steps their earlier steps do before the
  failed one runs again (an `enter` filling the same field counts).
- A value read before a rescue outranks the flow's empty declaration of
  it; "no products" and "0 products" (a header's empty cart) no longer
  read as a search that found nothing.
- A screen that fills most of a request is cut before splitting, so its
  questions are not asked one call each, and the budget is charged per
  part and framing.
- CI's clippy (Rust 1.99) rejects `assert!(….is_empty())` with no message.
@YellowSnnowmann YellowSnnowmann changed the title WIP: browser agent fixes from live tests on six sites WIP: browser agent fixes from live tests on ten utility sites Oct 6, 2026
@YellowSnnowmann
YellowSnnowmann marked this pull request as ready for review October 6, 2026 18:36
@YellowSnnowmann YellowSnnowmann changed the title WIP: browser agent fixes from live tests on ten utility sites feat: browser agent fixes from live tests on ten utility sites Oct 6, 2026
The per-file coverage gate failed on surface/tabs.rs (59%) and
surface/uncover.rs (87%): nothing exercised a navigation that timed out
on a drawn page, a press retried after a cover, or a link followed when
its press went nowhere. Three browser tests now drive each path through
the fake engine, including the cases that must keep the original reply
(nothing drawn yet, the page moved by itself). keep_in_tab loses two
unreachable early returns.
fresh_rows left out boxes text is set into (SetValue) but not boxes only
typed into (TypeText), which the rest of the flow treats as text boxes
too: one drawn with the list whose label mentions the text could be
offered and pressed as a suggestion. A unit test pins it.

The simulated ride form is now as strict as a page: a press takes only a
row on show for the box's current text, and a press anywhere else but the
box moves the focus on, closing the list and dropping unpicked text.
The approval prompt and the state line carry words from the pages a task
read, and a page chooses its own: an escape sequence in a button's name
could recolour, hide, or rewrite the prompt a person approves an
irreversible action from. Control characters and direction marks now
print as spaces.

follow's docs also say what the time limit does while a person answers:
they are never cut short, and the limit is checked on the next state.
It refuses bare vendor ids such as the engine's defaults
(anthropic/claude-sonnet-5) and takes openrouter/-prefixed ids from its
catalog, which task_live's default is; the OpenAI defaults in
.env.example apply on OpenRouter.
@senamakel
senamakel merged commit 6f5bb5b into tinyhumansai:main Oct 7, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants