← Selected work Dockline / Case study Source ↗

Dockline

Don't ask the model if the model was right.

Live · 40 eval cases · Updated 2026

Python · FastAPI · MCP · Evals

Hosted on Render. The first open may take up to about 60 seconds; the proof below remains available immediately.

Context
Independent agent-evaluation project · 2026
My role
Eval designer and sole engineer
Team
Solo build
Evidence
40 deterministic cases · complete traces

01 Problem

If the model grades the model, you will ship a confident wrong tool call. An ops agent on a warehouse API needs a judge that does not float.

02 Constraints

  1. 01The agent only acts through Clearbay MCP.
  2. 02Forty rule-scored cases. The model is never the judge.
  3. 03Tenant isolation, read-only writes, extra JSON fields, vague prompts must fail closed.
  4. 04Default router is deterministic. OpenAI tools are optional.

03 Architecture

04 Decisions

Eval harness is the product, not a later test.

Why Tenant leaks and extra fields are scoring rules, not vibes.

Tradeoff Cases are hand-written. That is cheaper than a wrong wave release.

Deterministic router by default.

Why You can demo isolation without an API key. The optional model is never the judge.

Tradeoff Vague language still needs a clarification case, not a guess.

Traces are first-class.

Why When a case fails, you read the tool call, not a chat bubble.

Tradeoff More logging than a toy agent. Necessary for an on-site loop.

05 Failure

  • Prompt tries another tenant.

    Isolation case fails the run. Tool never fires cross-tenant.

  • Read role attempts a write.

    Refusal case. MCP write tools require ROLE_OPS.

  • Extra JSON fields on a tool call.

    Schema case. Unknown fields rejected.

  • Vague prompt.

    Clarification case. No speculative write.

06 Object

07 Inspect

Live Open product ↗

Source GitHub ↗

Wake time Allow up to about 60 seconds on the first request.

08 What I would improve next

  1. Generate adversarial and property-based cases, then use mutation testing to prove the rules catch regressions.
  2. Build a replay corpus across model and prompt versions instead of treating one pass rate as permanent.
  3. Add latency and cost budgets plus explicit human escalation for requests that cannot be scored safely.