XY

Agents · Evaluation

The Model Is Never the Judge: Deterministic Agent Evals

Language models are useful actors and useful critics. They are the wrong final authority for tenant isolation, tool schemas, write permissions, and other invariants that must not drift.

Decision map Keep the actor flexible and the release gate explicit

The model or router may vary. The trace format and invariant scorers remain stable, and semantic review runs only after hard gates pass.

  1. 01 Prompt

    One tenant-bound request

  2. 02 Actor

    Deterministic router or optional model

  3. 03 MCP trace

    Tool, arguments, result, and final answer

  4. 04 Invariant gate

    Isolation, role, schema, disclosure

  5. 05 Review

    Optional model or human quality judgment

A fluent answer can still be a failed run

At JASCI I worked on an LLM action layer for warehouse operations. The difficult part was not getting a model to call a tool once. It was preserving the contracts that already protected a multi-tenant production system: identity comes from trusted context, writes are role-gated, schemas reject unknown fields, and repeated actions do not create repeated damage.

A model can produce an elegant explanation after choosing the wrong tenant or inventing an argument. Asking another model whether the result looks correct can be useful for language quality, but it does not turn a security invariant into a reliable test.

If a rule can be stated exactly, score it exactly before asking a model for an opinion.

Separate behavior from judgment

Dockline is an eval-first agent layer over Clearbay's warehouse MCP tools. A prompt enters a deterministic router by default, the router chooses a typed tool or clarification path, and every call produces a trace containing arguments, tool results, and final output. An optional model can replace the router, but it never replaces the scorer.

The suite contains 40 hand-authored cases across lookup behavior, tenant isolation, read-only write refusal, schema validation, clarification, and final summaries. Each case defines the input, allowed behavior, and explicit pass conditions. The scorer reads the trace and output as data.

prompt → router → MCP tool → trace → rule scorer
                    ↘ clarify
                    ↘ refuse

Turn production contracts into eval rules

The strongest evals come from rules the underlying system already needs. A read-only user must not execute a wave release. A tenant-scoped session must not retrieve another tenant's inventory. Tool arguments must match JSON Schema exactly. A vague request that could mutate state must ask for clarification instead of guessing.

Those assertions do not require semantic taste. The trace either contains a prohibited tool call or it does not. The arguments either include an unknown field or they do not. The final answer either leaks a protected location or it does not. A case can fail even when the lookup itself succeeded if the final response exposes information outside the allowed boundary.

  • Authorization: was a write attempted without the required role?
  • Isolation: did any tool argument or result cross the active tenant?
  • Schema: were unknown or malformed fields rejected before execution?
  • Uncertainty: did the agent clarify when required inputs were missing?
  • Disclosure: did the final answer reveal data the caller could not request?

Why an LLM judge is still useful—but second

Not every property is binary. Clarity, tone, completeness, and whether a summary preserves important nuance may benefit from model-based critique or human review. The mistake is letting that flexible judge decide the same gate as security and correctness.

I prefer a layered score: deterministic invariants first; task-specific structural checks second; optional semantic review third. If an invariant fails, the run fails regardless of how persuasive the explanation sounds. If all invariants pass, a model judge can help compare two acceptable summaries without being asked to certify authorization.

A trace is a debugging artifact, not telemetry decoration

A percentage alone cannot explain a regression. The useful unit is a replayable case with the prompt, routing decision, selected tool, arguments, result, refusal or clarification, and final text. That artifact lets an engineer distinguish a prompt problem from a schema change, service failure, or scorer bug.

This also makes model changes reviewable. A new model version may improve conversational quality while reducing clarification discipline. Without the same deterministic corpus and trace format, that tradeoff disappears inside an aggregate success rate.

What I would improve next

Forty cases are a credible beginning, not a permanent safety claim. I would add adversarial and property-based generation, then mutation-test the scorer to prove that removing a guard causes the expected cases to fail. Real incidents and near misses should become a versioned replay corpus.

I would also run the corpus across model, prompt, and tool-schema versions while tracking latency and cost budgets. Requests that cannot be evaluated safely should have an explicit human escalation path. The goal is not to prove that a model is universally good; it is to know exactly which behaviors the system is willing to ship.

Source notes

Claims you can inspect.

  1. Code score.py · deterministic rule scorer ↗

    The implemented tool, tenant-disclosure, result, and argument checks over replayable traces.

  2. Code cases.jsonl · versioned eval corpus ↗

    Forty hand-authored cases that separate expected behavior from the system solving them.

  3. Reference OpenAI Evals · evaluation framework ↗

    Primary documentation for datasets, evaluation logic, run logs, and repeatable model comparison.

  4. Reference MCP specification · authorization ↗

    Current protocol requirements for token audience, validation, and transport authorization boundaries.