Skip to content

feat(diagrams): the agent draws in the chat — render_diagram, phase 1 of the Rich Diagrams spec - #104

Open
ndemianc wants to merge 9 commits into
developfrom
feat/rich-diagrams
Open

ndemianc wants to merge 9 commits into
developfrom
feat/rich-diagrams

Conversation

@ndemianc

@ndemianc ndemianc commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

What this is

Ask the agent how something is put together and it answers with boxes and arrows typed out of characters. They break with the font, the theme and the width of the panel; they cannot be clicked or exported.

With this PR the agent has a tool, render_diagram. The model describes structure — nodes, edges, groups, one accent, never a coordinate or a colour. The editor validates the description, lays it out in one house style and paints it in the chat, in the editor's theme, with nodes that open the code they stand for.

This is phase 1 of the Rich Diagrams spec: Graph JSON only. Mermaid, Vega-Lite and raw SVG are not built.

The architecture fixture, light palette

The same fixture, dark palette

Two of the committed snapshots: what the SVG export writes. In the chat the colours are the editor's own theme tokens.

How it works

  1. A request from a client that can render carries the tool and a short block of rules.
  2. The model calls the tool. A placeholder holds the place, titled as soon as the title has streamed.
  3. The host makes the lossless fixes and validates. Valid: a record is stored and sent to the chat. Invalid: every error goes back to the model, once.
  4. The chat validates the record again, measures the text in the real font, lays it out and paints it from allow-lists.
  5. If the one repair pass does not produce a valid spec, the editor draws what is valid and says what it left out — or shows the errors, the source and Retry. Never an empty card.
Where What
diagram/schema.js validate.js repair.js The versioned schema, every error at once with a JSON Pointer, and the repair ladder
diagram/theme.js layout.js The style guide as numbers; a layered layout with orthogonal connectors
diagram/scene.js text.js ascii.js The painter (two allow-lists, textContent), the screen-reader outline, Mermaid and source export, the text fallback
diagram/tool.js service.js The tool and its rules; one conversation's diagrams and their one repair pass
diagram/links.js exportCheck.js stats.js bundle.js Links resolved inside the workspace; exports checked before they are saved; local counters; the shared modules inlined into the page
agent.js, providers/* The tool in the loop; a call's id and its streamed arguments
sessionEvents.js sessions.js extension.js Diagrams stored with the turn and replayed; stubs at compaction and get_diagram; link, export, retry
media/chat.html The card, its states, zoom, toolbar, fallback

The page's content security policy is unchanged. The shared modules are inlined under the nonce the page already has, and it still loads nothing.

Where it departs from the spec

The spec This PR Why
Graph JSON, Mermaid, Vega-Lite, opt-in raw SVG Graph JSON only Phase 1. The rules tell the model to use a table for numbers and a numbered list for sequences, not a format the chat would show as source
Layout by ELK.js An in-house layered layout ELK is EPL-2.0 in an otherwise MIT-clean repository, about 1.5 MB, and the extension has no build step. layout.layout() is the one entry point, so ELK can replace it
Validation by Ajv A small interpreter for the subset of JSON Schema the schema uses The same reasons
"Dedupe ids" and "drop edges to unknown nodes" are fixes made without a model Lossy fixes run only after the model's one repair pass has failed The spec's own table says semantic errors go to the model, "because intent is needed". An edge to billing when the node is bill is a typo; dropping it silently changes what the diagram says
The host accepts two actions from a diagram Four: openLink, export, retry, ascii The spec's own Retry button and text fallback need the other two. Each names a diagram by id; none carries a path or a URL
Rendering in a worker or sandboxed frame In the chat webview For Graph JSON the painter executes nothing the model wrote. Mermaid and Vega run third-party parsers over model text and should get the frame

The rule the ladder follows: an auto-fixed diagram never says something different from what the model wrote; a degraded one always says what it lost.

Also in here: a chat rendering bug, unrelated to diagrams

The first two commits fix a bug that is on develop today. An answer that mentioned a code fence in passing — the model quoted three backticks inside an inline span of four — was rendered as two one-character code blocks, and the rest of the paragraph came out one streamed fragment per line.

The renderer treated any three backticks, anywhere, as a fence. And once a "closing" fence had no newline after it, the streaming renderer froze every delta into its own block. A fence is now a line of its own, inline code can be delimited by any number of backticks, and only a block whose closing line is complete is frozen.

96399b6 pins what the renderer did before; 3b05fd3 is the fix. Replayed in Chromium with the session's exact text in 184 deltas: the paragraph was 57 lines and two code blocks before, and is one paragraph and none after. The two commits stand alone, and can be cherry-picked into their own PR if they should merge first.

The commits

Commit What
96399b6 Pins how the chat renders fences and inline code
3b05fd3 The chat fix
42159e5 The diagram schema, validator and repair ladder
fb2e67a The theme and the layout
f35fbc1 The painter, the text outline, the text fallback
ce02534 render_diagram in the agent, the host and the chat
822cfc2 The browser check, the real-editor check, the eval harness
db68f68 docs/RICH-DIAGRAMS.md, CLAUDE.md, the fixtures' README

Each one passes ./scripts/test-extensions.sh on its own. The gate was run on every committed tree in a separate checkout, not only on the tip.

Verification

  • The gate: 62 suite files, on macOS arm64 with Node 24, and on Linux (Ubuntu 22.04 arm64, Node 18) in a container with no network and a read-only checkout.
  • Twelve diagram suites, 251 tests. They include a 1,500-spec layout fuzz, 39 seeded broken specs with the exact outcome of each attempt, SVG snapshots in both palettes, the gallery at chat-panel widths, the real runAgent loop against a scripted provider, and extension.js's own functions sliced out and run against a stand-in for vscode.
  • The chat page in a browser (scripts/diagram-browser-check.js): the shipped chat.html in headless Chrome under the real policy, 155 checks across both themes. Hostile labels are on the page as text and nothing ran; no policy violation; no request.
  • The real editor (scripts/diagram-editor-check.js): a second, throwaway instance of the dev build loads this branch's extension and talks to a stand-in provider on localhost, 17 checks. The tool and its rules reach the model, the diagram is painted in the real webview in the editor's colours, a linked node opens its file at the symbol, and the session on disk holds the diagram.
  • Mutation testing: 168 single-edit defects seeded across the modules, the host glue, the page, the eval harness and the chat fix. Twelve survived at first: nine were gaps in the suites, now closed, and three were redundant code, now removed. All 165 that still apply are caught.
  • Timing: the diagram suites and the chat suite also pass with every timer delayed by 15 ms.

Not verified

  • No live model has drawn a diagram. Every "model" above is a script. scripts/diagram-eval.js is the spec's eval — 30 prompts that should produce a diagram and 10 that should not, run through the real agent loop — and it has not been run, because --run makes billed calls. Nothing here says how often a real model's first attempt is valid, or whether it reaches for the tool at all.
  • CI's own runner (ubuntu-latest, x64, Node 24). This PR is its first run there.

For the reviewer to decide

  • On by default. levelcode.ai.diagrams.enabled defaults to true, and costs about 970 tokens of tool and rules on every agent request while it is on. It is constant for a session, so it sits in the cached prefix. Flip the default if this should ship dark until the eval has run.
  • A model can be switched off with diagrams: false in its providers/catalog.js row. That is the lever for one that fails the eval.
  • docs/RICH-DIAGRAMS.md is an implementation record, not the spec. The spec is a private document.

Known limits

Measured over 6,000 random specs at the model-facing limits. None of the nine gallery diagrams shows any of the three, at its natural width or at 560, 420 and 320 px.

Limit How often
An edge label falls back to a halo over a line 6.5% of specs, mostly dense fan-ins in a top-to-bottom flow
Two connectors swap lanes in one gap 0.3%
A group's name has a connector passing behind it 2.7% of specs with groups

Each is bounded in the layout suite, so it cannot get worse quietly. The rest are in the doc.

To try it

On this branch, in the checkout that has vscode/, ./scripts/run-dev.sh is enough. From a git worktree, which has no vscode/ of its own, load the worktree's extension into the build that exists:

./scripts/run-dev.sh --extensionDevelopmentPath=/absolute/path/to/worktree/extensions/levelcode-ai

The title bar says [Extension Development Host]. In agent mode, ask something whose answer is structure: "how does a request flow through this app?"

…fore changing it

The chat's Markdown renderer lives inline in media/chat.html, and no suite ran
it. The next commit changes how it finds a code fence, so this one writes down
what it does today, against today's code:

  an ordinary fenced block      a code block; the language line is not in it
  two blocks, prose between     and a block still being typed is shown as code
  inline code                   escaped, never formatted; a lone backtick is text
  a path in inline code         a file chip, when it names a real project file
  a fence inside a list item    still a code block, indented as it was
  a fence stuck to a line       "Run this: ```bash" opens a block, and a closing
                                fence stuck to the last line of code closes one
                                — sloppy, and already tolerated

The functions are sliced out of chat.html itself, the way shHighlight.test.js
does it, so the suite runs the shipped code and not a copy of it.
…aragraph came out a fragment per line

An answer that mentioned a code fence in passing — "each closed by a plain
```` ``` ````" — was rendered as two code blocks holding one backtick each, and
the rest of the paragraph arrived one streamed fragment per line: "gr",
"ounded ent", "irely in the actual code". Reported from a real session. The
model had quoted the fence correctly, in an inline span of four backticks.

Two things were wrong, and the second is what made it look so bad.

render() split the text on every run of three backticks, wherever it stood. So
```` ``` ```` was three fences: a block containing "` ", and then an unfinished
block containing the rest of the message.

lastStableIndex() decides what the streaming renderer may freeze — move out of
the live tail and into the page for good. It counted fences the same way, and
when the text after an even-numbered one had no newline yet, it answered "all
of it". From then on every delta was frozen as it arrived, each in its own
<div>.

A fence is a line. mdSegments() reads the text line by line: a block opens on a
line that is, after its indentation, three or more backticks and a language (or
anything else without a backtick in it), and closes on a line ending in a run
at least as long as the one that opened it. render() and lastStableIndex() both
use it, so they cannot disagree again. A block counts as finished only once its
closing line has its newline; text after it stays live until the next block
finishes.

Inline code takes the count seriously too: a run of N backticks opens a span
that the next run of exactly N closes, on the same line. That is how three
backticks are quoted (inside four), and one (inside two).

Kept, because models do it and the old code let them: a fence stuck to the end
of a line of prose ("Run this: ```bash") still opens a block — when nothing but
a one-word language follows it and it is not the closing half of an inline
span — and a closing fence stuck to the last line of code still closes one.

Not changed: tildes are not fences; and a line of prose that ENDS in three bare
backticks still opens a block, as it did. CommonMark would not; a model that
wants to say "```" has an inline span for it, and now it works.

chatMarkdownFences.test.js gains the report itself, whole and streamed — in
sixty different chunkings, and one character at a time; the property that
makes streaming safe (streamed and whole are the same page, and only finished
blocks are frozen) over six documents and 400 random mixes of backticks,
newlines and words; N-backtick spans; fences nested by length; language lines
that say more than a word; and that nothing in a message becomes markup. The
six pins from the last commit pass unchanged. The ten new tests fail on the old
code.

Replayed in Chromium against the shipped page with the session's exact text, in
184 deltas: the paragraph was 57 lines and two code blocks before, and is one
paragraph and none after. 21 single-edit mutations of the new code are each
caught by the suite.
…rror at once, and the repair ladder

First slice of rich diagrams; docs/RICH-DIAGRAMS.md arrives with the last one.
The agent draws flows and architectures out of box-drawing characters. The plan
is that it describes STRUCTURE instead — nodes, edges, groups, one accent,
never a coordinate or a colour — and the editor draws. This commit is the part
that decides whether a description may be drawn. Nothing calls it yet.

  diagram/schema.js     the versioned schema, and a small interpreter for the
                        subset of JSON Schema it uses. One object is both the
                        tool's input schema and what a spec is checked against.
                        Two tiers of limits: HOUSE is what a model is held to
                        (12 nodes; labels 28, second lines 32, edge labels 20,
                        title 80; groups two deep), HARD is what the renderer
                        will still draw (24 nodes, 48 edges).
  diagram/validate.js   schema, then semantics: edge ends exist, ids are unique,
                        one accent, groups known, acyclic, at most two deep.
                        EVERY error, each with a JSON Pointer, what was
                        expected and the valid options:
                          /edges/1/to: unknown node "billing". Known ids: in, jev, bill, rev.
  diagram/repair.js     the ladder. prepare() takes one call up it.

The ladder departs from the spec in one place, on purpose. The spec lists
"dedupe ids" and "drop edges to unknown nodes" among the fixes made without a
model, while its own table says semantic errors go to the model, "because
intent is needed". Both cannot hold: an edge to "billing" when the node is
"bill" is a typo the model fixes in one pass, and dropping it silently changes
what the diagram says. So:

  1  auto-fix, lossless only   lenient JSON (comments, trailing commas, smart
                               quotes, bare keys), synonyms ("diamond" for
                               decision, source/target for from/to), slugged
                               ids, an edge that names a node by its label,
                               long labels shortened with the full text kept
  2  every error, once         returned for the model's one repair pass
  3  degrade                   only now the lossy fixes: edges to nowhere
                               dropped, duplicates renamed, one accent kept,
                               counts up to the HARD tier — each loss named

An auto-fixed diagram never says something different from what the model
wrote; a degraded one always says what it lost.

accept() is the same gate for a spec that is read back rather than received —
from a session file, say: validate, hand on only the fields the schema
declares, and fall back to the rungs that need no model.

No Ajv: it is a dependency, and this extension has neither dependencies nor a
build step.

test/fixtures/diagrams/corpus.json holds 39 broken specs with the exact outcome
of each attempt. They are seeded from the ways models are known to get this
wrong, not collected in the field; real ones should be added as they turn up.

diagramSchema (18 tests) and diagramRepair (65).
…house style

  diagram/theme.js    the style guide as numbers: title 15, node name 13
                      semibold, second lines and edge labels 11.5, nothing
                      under 10.5; radius 8, border 1.25, padding 12; groups
                      inset 16; the accent a low-opacity fill and a 2px border.
                      Colours are editor theme tokens, with two fixed palettes
                      for exported files.
  diagram/layout.js   layout(spec, opts) gives geometry; inspect(geometry) lists
                      whatever overlaps, overflows or leaves the frame.

The spec names ELK.js. This is not ELK: ELK is EPL-2.0 in a repository that is
otherwise MIT-clean, about 1.5 MB, and this extension has no build step. The
spec's own open question asks whether a simpler layered layout is needed as a
fallback; this is that layout, as the only one. layout() is the single entry
point, so ELK can replace it without touching anything else.

What it does: ranks by longest path; crossings reduced by sweeps; positions
across the flow solved as constraints and then lined up by medians; a port per
edge; a track per connector in each gap, so lines cross but never run along
each other; and each edge label placed BESIDE its line, scored against nodes,
lines and other labels. A flow asked for left-to-right that will not fit its
column is laid out top-to-bottom instead.

A group's name sits in the top-left of its frame — which, when the flow runs
down, is exactly where lines come in. No connector runs through it. The name
slides along the frame to the nearest clear stretch; if there is none and the
lines have under 72px to move, they are moved to the far side of the name (the
frame grows by that much, never past the column); otherwise the name stays and
is marked to be drawn over the line.

Measured over 6,000 random specs at the model-facing limits: no node overlaps
another, no connector crosses a node or an unbacked group name, nothing leaves
the frame. Three known limits, each bounded in the suite so it cannot quietly
get worse:

  an edge label that falls back to a halo over a line   6.5% of specs
  two connectors that swap lanes in one gap             0.3%
  a group's name with a connector behind it             2.7% of specs with groups

None of the nine gallery diagrams (fixtures/diagrams/gallery.json) shows any of
the three, at its natural width or at 560, 420 and 320px. The twelve-node one
lays out in under a millisecond once warm; the slowest of the 6,000 took 7ms.

diagramLayout (23 tests, one of them a 1,500-spec fuzz).
  diagram/scene.js   geometry to a tree of elements, built from two allow-lists:
                     seven tags (svg g rect path text title style) and a set of
                     attributes with no href, no style, no id and no on*.
                     mount() turns it into DOM with createElementNS and
                     textContent; toSvg() into an escaped string for a file. A
                     label is never parsed as markup by either, and both refuse
                     anything off the lists rather than trust whoever built the
                     tree.
  diagram/text.js    the outline a screen reader is given (nodes, then edges, in
                     reading order); the one-line stub that stands in for a spec
                     once a conversation is compacted; Mermaid and source
                     export.
  diagram/ascii.js   for when a picture cannot be shown: the SAME layout on a
                     character grid, so it puts things where the picture would.
                     Plain ASCII; East Asian wide characters and emoji count as
                     two cells, combining marks as none.

A linked node carries data-lc-link = its NODE ID, never the path: whoever
handles the click looks the link up in its own copy of the spec.

A group's name that the layout could not keep clear of a connector is drawn
after the connectors, with the halo an edge label gets, so the line reads as
passing behind the word.

Snapshots of three gallery diagrams in both palettes (fixtures/diagrams/
snapshots): a change to the style, the layout or the painter shows up there as
a diff. diagramScene (23 tests) and diagramText (17).
Ask the agent how something is put together and it answered with boxes and
arrows made of characters, which break with the font, the theme and the width
of the panel. Now it has a tool. It describes the structure; the editor
validates it, lays it out and paints it in the chat, in the editor's theme,
with nodes that open the code they stand for.

Graph JSON only: phase 1 of the spec. Mermaid, Vega-Lite and raw SVG are not
built, and the rules tell the model to use a table for numbers and a numbered
list for sequences rather than a format the chat would show as source.

The model's side (diagram/tool.js)
  render_diagram   its input schema is the validator's own schema object
  the rules        when a diagram earns its place; never draw with characters;
                   asked for a diagram, draw it here rather than writing a file
                   of diagram source; the title states the takeaway; one
                   accent; split above twelve nodes; link nodes to workspace
                   files. One worked example, itself a valid spec.
  the answer       {"ok":true,"id":"d-1"}, or every error at once
  About 520 tokens of tool and 450 of rules on every request from a client
  that can render — constant for a session, so they sit in the cached prefix.

One conversation's diagrams (diagram/service.js)
  A call that fails validation is sent back ONCE, with every error. A second
  failure is drawn degraded, with a banner naming each loss and a Retry button
  that the user presses, not the editor. Three bounces in one run and every
  later call is final. A diagram still owed its repair when the run ends — the
  model gave up, ran out of steps, was stopped — is settled from its last
  attempt before agentDone, so no placeholder is left waiting. Arguments cut
  off at the token limit are asked for again, never repaired.

In the chat (media/chat.html)
  A placeholder from the moment the call starts, titled as soon as the title
  has streamed; then the picture. The shared modules are inlined into the page
  by diagram/bundle.js under the page's existing nonce: the content security
  policy is unchanged, and the page still loads nothing. Text is measured in
  the real font. A column too narrow for a left-to-right flow turns it
  downward, and it turns back when the column widens. A full-size view with
  zoom and pan. The "auto-fixed" badge and its details; the degraded banner;
  the errors and the source when nothing can be drawn. Copy source, SVG, PNG,
  open as Mermaid, insert into a Markdown file. The title is the accessible
  name, and a text outline is there for screen readers. If painting fails, or
  the modules never loaded, the same diagram is shown as text — drawn by the
  page in the first case and by the host in the second.

What the page may ask of the editor
  openLink, export, retry and ascii. Each names a diagram by id and is looked
  up in the host's own records. A click sends the NODE id; the path comes from
  the record and is resolved again at that moment (diagram/links.js): inside a
  workspace folder once ".." and symlinks are followed, and a file. SVG and PNG
  come back from the page, so they are checked before they are written
  (diagram/exportCheck.js). The spec lists two actions; its own Retry button
  and text fallback need the other two.

Sessions and context
  The final record — spec, status, fixes, what was lost — is stored with the
  turn and replayed when a chat is reopened; nothing is repaired again. A
  session file is input too: a stored spec passes the validator on its way
  back in, in the host and in the page. At compaction a diagram becomes one
  line ("diagram: <title>, 4 nodes, id d-17"), and get_diagram is offered from
  then on. A session export writes each diagram as a fenced Mermaid block.

Who gets it
  A run carries client.render: rich or ascii. Rich is the chat webview, unless
  levelcode.ai.diagrams.enabled is off or the model's catalog row says
  diagrams: false — the switch for a model that fails the eval. An ascii client
  is offered neither the tool nor the rules. Chat mode has no tools and is
  unchanged.

Counters (diagram/stats.js, "AI: Diagram Statistics")
  First-pass valid rate, fixes by rung, error classes per model, render time,
  tokens per diagram, and answers that drew with characters anyway. Kept in
  the editor's own storage and sent nowhere; enums and numbers only, never a
  label, a title or a path.

Providers: onToolStart now carries the call's id, a new onToolInput streams its
arguments, and a turn returns the raw text of arguments that did not parse.
All additive.

Four existing suites change. agentNoWorkspace and sessionsUi pinned the shape
of the tool list in source, and now pin the new shape. authRetryCallers and
sessionExpiredHost slice compactAgentMemory and resumeSession out of
extension.js, and needed the names those now use.

Tests: diagramAgent (22 — the real runAgent with a scripted provider),
diagramSession (12), diagramLinks (9), diagramHost (27 — extension.js's own
functions against a stand-in for vscode), diagramStats (8), diagramUi (11).
…, and an eval for models

Three scripts for the three things the suites cannot show. None is in the
gate, which is plain Node and stays that way.

scripts/diagram-browser-check.js — the shipped chat.html in headless Chrome,
built the way the host builds it (the same bundle, the same policy) and fed
records made by the real service. 155 checks across both themes: every diagram
is an SVG of the painter's elements only; hostile labels are on the page as
text and nothing ran; no policy violation and no request; the placeholder, the
badge, the banner, the failed view; a click names a node, never a path; SVG
and PNG export pass the host's own check; and a painter that throws, or a page
with no diagram modules at all, both end in the text fallback. Each step waits
for its result, not for a length of time: the first version waited, and failed
one run in three.

scripts/diagram-editor-check.js — the real editor. A second, throwaway instance
of the dev build (its own profile, sessions folder and workspace) loads THIS
checkout's extension in place of the built-in one, talks to a stand-in provider
on localhost, and is driven through the DevTools protocol. 17 checks: the tool
and its rules reach the model; the diagram is painted in the real webview in
the editor's colours; a wider column re-lays it out; a linked node opens its
file at the symbol; the session on disk holds the diagram; and the answer
around it, which quotes a code fence, is one paragraph (the chat fix earlier in
this branch).

It exists because of a mistake worth writing down. The feature was first called
done having only ever run in a browser, with instructions for trying it that
could not have worked: the work sat uncommitted in a git worktree, and a
worktree has no vscode/ — run-dev.sh runs the extensions of the checkout that
does. The editor that was then tried was develop's, with no diagrams in it.
What does work, by hand, from the checkout that has vscode/:

  ./scripts/run-dev.sh --extensionDevelopmentPath=<checkout>/extensions/levelcode-ai

scripts/diagram-eval.js — the spec's eval: 30 prompts that should produce a
diagram and 10 that should not (fixtures/diagrams/eval-prompts.json), each run
through the real agent loop and scored on format choice, first-pass validity,
fixes by rung, degraded rate, error classes, tokens, node-count overruns, title
quality and answers that drew with characters. --dry-run uses a scripted model
and no network. --run makes billed calls on your key, says how many before the
first one, and is refused without a model and a key; the workspace is an empty
temporary folder and every approval is answered no.

It has not been run against any model. That is the next thing this feature
needs: nothing here says how often a real model's first attempt is valid, or
whether it reaches for the tool at all.

diagramEval (16 tests) holds the harness to its own arithmetic: a scripted
model whose mistakes are known comes out at exactly the numbers its script
implies, and nothing is sent without --run, a model and a key.
…ow to check it

docs/RICH-DIAGRAMS.md is the implementation record for the Rich Diagrams spec
(2026-10-04), under the spec's own section names, because the code cites them.
Per section: the rule, what the code does about it, and each place the build
chose differently, with the reason, so the choice can be argued again. A status
per requirement; the known limits with their measured rates; what is verified
and what is not — no live model has drawn a diagram yet, and the eval has not
been run.

It is not the spec. The spec is a private document; its links, and one line
about pricing, are left out of a public repository.

CLAUDE.md gains the feature's entry and three conventions that cost time to
learn: the shared diagram modules are pasted into a script block and must never
spell a script tag or an HTML comment opener; the host suites' brace matcher
cannot read a backtick inside a regex literal; and work in a worktree is not in
the editor until it is loaded with --extensionDevelopmentPath.

test/fixtures/diagrams/README.md says what each fixture is for, and how to add
a real broken spec to the corpus.
Copilot AI balanced review requested due to automatic review settings October 5, 2026 04:38
Comment thread extensions/levelcode-ai/diagram/layout.js Fixed
Comment thread extensions/levelcode-ai/diagram/theme.js Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Untrusted schema versions can crash validation, and several repair, replay, compaction, statistics, and accessibility paths remain incorrect.

Review effort: Balanced
Findings: 2 High severity · 5 Medium severity

Open (7)
What changed in this PR

Adds phase-one rich Graph JSON diagrams to agent chat, including validation, rendering, persistence, export, accessibility fallbacks, and evaluation tooling. It also fixes streamed Markdown fence handling.

Changes:

  • Adds the render_diagram agent tool and provider streaming support.
  • Implements safe diagram layout, rendering, links, exports, persistence, and statistics.
  • Adds extensive unit, browser, editor, snapshot, and evaluation coverage.
File Description
CLAUDE.md Documents diagram architecture and conventions.
docs/​RICH-DIAGRAMS.md Records the feature design.
extensions/​levelcode-ai/​agent.js Integrates diagram tools into agent runs.
extensions/​levelcode-ai/​diagram/​ascii.js Adds text fallback rendering.
extensions/​levelcode-ai/​diagram/​bundle.js Bundles shared modules into the webview.
extensions/​levelcode-ai/​diagram/​exportCheck.js Validates exported SVG and PNG data.
extensions/​levelcode-ai/​diagram/​layout.js Implements layered diagram layout.
extensions/​levelcode-ai/​diagram/​links.js Resolves safe workspace links.
extensions/​levelcode-ai/​diagram/​repair.js Implements normalization and repair.
extensions/​levelcode-ai/​diagram/​scene.js Builds allow-listed SVG scenes.
extensions/​levelcode-ai/​diagram/​schema.js Defines the versioned Graph JSON schema.
extensions/​levelcode-ai/​diagram/​service.js Manages diagram lifecycle and records.
extensions/​levelcode-ai/​diagram/​stats.js Collects local rollout metrics.
extensions/​levelcode-ai/​diagram/​text.js Generates outlines, stubs, and Mermaid.
extensions/​levelcode-ai/​diagram/​theme.js Defines diagram styling and theme tokens.
extensions/​levelcode-ai/​diagram/​tool.js Defines tools and model instructions.
extensions/​levelcode-ai/​diagram/​validate.js Validates schema and semantics.
extensions/​levelcode-ai/​extension.js Wires host actions, persistence, and exports.
extensions/​levelcode-ai/​media/​chat.html Adds diagram cards and zoom UI.
extensions/​levelcode-ai/​package.json Registers settings and statistics command.
extensions/​levelcode-ai/​providers/​anthropic.js Streams tool IDs and partial arguments.
extensions/​levelcode-ai/​providers/​catalog.js Adds per-model diagram capability.
extensions/​levelcode-ai/​providers/​index.js Forwards diagram streaming callbacks.
extensions/​levelcode-ai/​providers/​openaiCompat.js Handles streamed OpenAI tool arguments.
extensions/​levelcode-ai/​providers/​translate.js Preserves malformed raw tool arguments.
extensions/​levelcode-ai/​scripts/​diagram-browser-check.js Adds browser-level rendering checks.
extensions/​levelcode-ai/​scripts/​diagram-editor-check.js Adds real-editor integration checks.
extensions/​levelcode-ai/​scripts/​diagram-eval.js Adds model evaluation harness.
extensions/​levelcode-ai/​sessionEvents.js Stores and exports diagram events.
extensions/​levelcode-ai/​sessions.js Persists diagrams with sessions.
extensions/​levelcode-ai/​test/​agentNoWorkspace.test.js Updates host-gated tool assertions.
extensions/​levelcode-ai/​test/​authRetryCallers.test.js Stubs diagram compaction behavior.
extensions/​levelcode-ai/​test/​chatMarkdownFences.test.js Tests streamed Markdown fence handling.
extensions/​levelcode-ai/​test/​diagramAgent.test.js Tests agent diagram integration.
extensions/​levelcode-ai/​test/​diagramEval.test.js Tests evaluation logic.
extensions/​levelcode-ai/​test/​diagramHost.test.js Tests host wiring and security.
extensions/​levelcode-ai/​test/​diagramLayout.test.js Tests layout and fuzz cases.
extensions/​levelcode-ai/​test/​diagramLinks.test.js Tests workspace-link containment.
extensions/​levelcode-ai/​test/​diagramRepair.test.js Tests the repair ladder.
extensions/​levelcode-ai/​test/​diagramScene.test.js Tests SVG safety and snapshots.
extensions/​levelcode-ai/​test/​diagramSchema.test.js Tests schema validation.
extensions/​levelcode-ai/​test/​diagramSession.test.js Tests persistence and replay.
extensions/​levelcode-ai/​test/​diagramStats.test.js Tests metric aggregation and privacy.
extensions/​levelcode-ai/​test/​diagramText.test.js Tests textual representations.
extensions/​levelcode-ai/​test/​diagramUi.test.js Tests webview integration and safety.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​README.md Documents diagram fixtures.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​corpus.json Adds malformed-spec corpus.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​eval-prompts.json Adds model evaluation prompts.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​gallery.json Adds representative diagrams.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​arch.dark.svg Adds dark architecture snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​arch.light.svg Adds light architecture snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​decision.dark.svg Adds dark decision snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​decision.light.svg Adds light decision snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​jev.dark.svg Adds dark routing snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​jev.light.svg Adds light routing snapshot.
extensions/​levelcode-ai/​test/​sessionExpiredHost.test.js Adds diagram service test stub.
extensions/​levelcode-ai/​test/​sessionsUi.test.js Updates session integration assertions.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +51 to +54
const model = String(ev.model || 'unknown').slice(0, 80);
if (!s.byModel[model] && Object.keys(s.byModel).length >= MAX_MODELS) { return s; }
const m = s.byModel[model] || (s.byModel[model] = { calls: 0, outcomes: {}, errors: {} });
m.calls++;
return { ok: false, errors };
}
const v = spec.v === undefined ? schema.VERSION : spec.v;
const known = schema.SCHEMAS[v];
Comment on lines +1037 to +1041
if (tu.name === diagramTool.RENDER_DIAGRAM.name && rich) {
const rawArgs = turn.raw && turn.raw.get(tu.id);
const parsed = (turn.stop_reason !== 'max_tokens' && typeof rawArgs === 'string') ? diagramRepair.parseLenient(rawArgs) : null;
const out = (parsed && parsed.ok) ? ctx.diagrams.render(rawArgs, { key: tu.id, model: ctx.model }) : ctx.diagrams.truncated({ key: tu.id, model: ctx.model });
for (const m of out.post) { ctx.post(m); }
Comment on lines +249 to +260
function stubsFor(messages) {
const out = [];
for (const msg of (Array.isArray(messages) ? messages : [])) {
if (!msg || msg.role !== 'assistant' || !Array.isArray(msg.content)) { continue; }
for (const b of msg.content) {
if (!b || b.type !== 'tool_use' || b.name !== tool.RENDER_DIAGRAM.name) { continue; }
const r = byToolUse(b.id);
if (r && r.spec) { out.push(text.stub(r)); }
}
}
return out;
}

<div id="imgzoom" hidden role="dialog" aria-modal="true" aria-label="Attached image"><img alt=""></div>

<div id="lcdZoom" hidden role="dialog" aria-modal="true" aria-label="Diagram, full size">
const descId = 'lcd-desc-' + (++lcdSeq);
svg.setAttribute('aria-describedby', descId);
stage.appendChild(svg);
stage.tabIndex = 0; stage.setAttribute('role', 'button'); stage.setAttribute('aria-label', 'Open the diagram full size'); stage.title = 'Click to open full size';
Comment on lines 143 to +149
const text = messageText(m.content);
if (!text.trim()) { continue; } // e.g. an assistant turn that was only tool calls
out.push({ role: m.role, text });
if (text.trim()) { out.push({ role: m.role, text }); } // (an assistant turn that was only tool calls has none)
if (m.role === 'assistant' && byKey.size && Array.isArray(m.content)) {
for (const b of m.content) {
const r = b && b.type === 'tool_use' && b.name === DIAGRAM_TOOL ? byKey.get(String(b.id)) : null;
if (r && !replaced.has(String(r.id))) { out.push({ role: 'diagram', key: String(r.key), record: r }); }
}
…eant to be; a dead store goes

Two findings from the code-quality review. Both are right.

theme.js, charEm() — "the condition c > 0xffff is always false". The width
estimate used where there is no canvas asked "CJK and other full-width scripts"
(U+2E80 and up, 1.0 em) before "emoji / astral" (above U+FFFF, 1.1 em). Every
code point above U+FFFF is above U+2E80 too, so the second question was never
reached and an emoji was measured at 1.0 em. The estimate is meant to err wide
— the failure that matters is text overflowing its box — so the questions are
now asked the other way round.

Nothing on screen changes. The chat measures with the real font, and no fixture
contains a character beyond the basic plane, so no snapshot moves. What changes
is the fallback when the canvas refuses the font, and the estimates made in
Node.

layout.js — "the value assigned to span here is unused". When several
connectors share a side too short for them, the node grows to hold exactly what
they need; span was recomputed afterwards and never read. The recompute is
gone and span is a const. it.cs, set on the same line, IS read later, and
stays. 6,000 random layouts and the nine gallery diagrams are byte-identical to
the code before.

diagramLayout gains a test for the estimate's classes: narrow, ordinary and
wide letters, digits and capitals, a CJK character at a full em, an emoji at
1.1 em and counted as one character. It fails on the old order, on exactly the
condition the review named.

62 suite files pass on macOS and in a Linux container.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants