Skip to content

feat(evals): add LLM-as-judge scoring - #8515

Closed
sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-llm-judge
Closed

sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-llm-judge

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Substring checks measure phrasing, not correctness — every live failure so far
was a valid paraphrase or an over-specific assertion. This adds a judge model
that scores an answer against a weighted rubric and returns structured numbers,
so the suite can assert behavior instead of wording.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8514

What changed

  • judge.ts — rubric + prompt, judgeAnswer, JSON parsing/clamping, verdict
  • judge.test.ts — key-free coverage: parsing, fenced JSON, clamping, weights,
    missing criteria, non-JSON
  • judge.live.test.ts — opt-in: a grounded answer outscores an invented one
  • harness.ts — runScenario accepts an optional judge and adds a judge
    check; deterministic checks are unchanged
  • test:evals:judge script; README documents it

How it works

const verdict = await judgeAnswer({
  completion, model: 'deepseek-chat',
  userMessage, answer,
  evidence: JSON.stringify(toolResults),
  rubric: { criteria: [
    { id: 'grounding', description: 'every claim is supported by the evidence' },
    { id: 'completeness', description: 'answers the user request' },
  ], minScore: 0.7 },
})
// => { scores, rationale, weightedScore, passed }

Because the judge transport is an ordinary OpenAI-compatible completion, it can
be recorded and replayed deterministically with the record/replay layer.

Test plan

  • judge.test.ts → 7/7 (parsing, weights, clamping, failures)
  • bun run test:evals → 19/19
  • bun run check:test-patterns passes
  • judge.ts and the harness change type-check against the real loop signature
  • bun run test:evals:judge live run (grounded > invented)
  • Full bun run type-check — run in CI

Follow-up

Wire the judge into selected scenarios (replace brittle finalContent regexes
where a rubric is more honest) once the live judge run confirms the prompt.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Substring checks measure phrasing, not correctness. judgeAnswer scores an
answer against a weighted rubric with a judge model and returns structured
scores; runScenario gains an optional judge that adds a judge check. The
judge transport is an injectable OpenAI-compatible completion, so a recorded
transcript can replay it deterministically.

- judge.ts: rubric, prompt, JSON parsing/clamping, verdict
- judge.test.ts: parsing/weighting/clamping (key-free)
- judge.live.test.ts: grounded answer outscores an invented one (opt-in)
- test:evals:judge script; README documents it
@sudoKrishna
sudoKrishna requested a review from a team as a code owner October 1, 2026 14:09
@vercel

vercel Bot commented Oct 1, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 2/5

[Medium risk] Adds evaluation framework for agent tool-use behavior.

The PR is not ready to merge because the judge lacks grounding evidence and the scripted and live harnesses can report misleading results.

Findings

  1. P1 Judge lacks tool results ▶
  2. P1 Scripted turns ignore feedback ▶
  3. P1 Live calls receive fabricated results ▶
  4. P2 Live checks reject paraphrases ▶
  5. P2 Invalid rubrics skew verdicts ▶
  6. P2 Relative imports violate app convention ▶

Summary

The PR adds scripted and opt-in live agent-tool-use evaluations, an executor harness, report generation, and an injectable LLM judge.

  • The judge lacks tool-result evidence for grounding, while scripted scenarios do not verify result feedback.
  • Live fixture handling and wording checks can distort pass rates.
  • Rubric validation and app import conventions need attention.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  S[Scenario] --> H[Tool-loop harness]
  H --> T[Mocked tool results]
  T --> L[Streaming tool loop]
  L --> A[Final answer]
  A --> D[Deterministic checks]
  A --> J[Optional judge]
  H --> J
  D --> R[Eval report]
  J --> R
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add LLM-as-judge scoring"

rubric: options.judge.rubric,
userMessage: scenario.userMessage,
answer: finalContent,
evidence: JSON.stringify(invocations),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Judge lacks tool results When a scenario enables the judge, it receives tool names, arguments, and success flags, but not the returned outputs. If the answer says the rate limit is 100 requests per minute, the judge cannot see the retrieved rate limit and cannot reliably score whether that claim is grounded. Pass the tool results as evidence.

Comment on lines +131 to +145
function createScriptedCompletion(scenario: AgentToolUseScenario): OpenAICompatCreateCompletion {
let turnIndex = 0
return async () => {
const turn = scenario.script[turnIndex]
turnIndex += 1
if (!turn) {
throw new Error(
`Scenario "${scenario.id}" requested model turn ${turnIndex} but only ${scenario.script.length} are scripted`
)
}
return (async function* () {
for (const next of turnToChunks(turn)) yield next
})()
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Scripted turns ignore feedback The scripted completion advances by turn number without reading the messages sent to it. If the loop stops feeding back a tool result or error, the script still emits its planned retry or answer, so retrieval and recovery scenarios can pass despite that regression. Check the tool messages at the completion boundary.

toolsMockFns.mockExecuteTool.mockImplementation(
async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => {
const startedAt = Date.now()
const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Live calls receive fabricated results Live tool calls consume results queued for the scripted calls, without matching the model’s actual arguments. If the model makes an extra retry, the exhausted queue returns a successful empty result. That fabricated result reaches the model, so its answer and the reported pass rate no longer reflect the intended tool evidence. Match results to live calls or fail when a fixture runs out.

Comment on lines +248 to +255
expect: {
toolCallSequence: ['get_weather', 'get_news'],
finalContent: /12°C[\s\S]*transit strike ends/i,
maxIterations: 3,
successfulToolCalls: 2,
},
/** The two tools are independent; a real model may emit them in either order. */
liveExpect: { toolCallSequence: undefined, requiredTools: ['get_weather', 'get_news'] },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Live checks reject paraphrases The live expectations keep this exact weather-and-news answer pattern. A correct answer such as “12 degrees Celsius” with a paraphrased headline fails the check, making the pass rate depend on wording rather than correctness. The weather scenario likewise retains an exact 17°C check; relax or judge these live answers.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +118 to +129
const weight = criterion.weight ?? 1
scores[criterion.id] = score
weighted += score * weight
totalWeight += weight
}

const weightedScore = totalWeight === 0 ? 0 : weighted / totalWeight
return {
scores,
rationale: typeof parsed.rationale === 'string' ? parsed.rationale : '',
weightedScore,
passed: weightedScore >= (rubric.minScore ?? 0.5),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Invalid rubrics skew verdicts Weights and minScore are not range-checked. A negative weight can put the weighted score outside the stated 0–1 range, while a negative threshold can pass a zero-scored answer. Validate positive finite weights and a threshold within 0–1 to avoid misleading verdicts.

import { adaptOpenAIChatToolSchema } from '@/providers/tool-schema-adapter'
import type { ProviderToolConfig, TimeSegment } from '@/providers/types'
import type { ToolResponse } from '@/tools/types'
import { type JudgeRubric, judgeAnswer } from './judge'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate app convention This module imports ./judge, although the repository requires absolute @/... imports in apps/sim. The same pattern appears in report.ts and executor-harness.ts. These imports need to use the app alias before merging to satisfy that requirement.

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@sudoKrishna sudoKrishna closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): add LLM-as-judge scoring

1 participant