Dockline
Don't ask the model if the model was right.
Live · 40 eval cases · Updated 2026
Python · FastAPI · MCP · Evals
Hosted on Render. The first open may take up to about 60 seconds; the proof below remains available immediately.
- Context
- Independent agent-evaluation project · 2026
- My role
- Eval designer and sole engineer
- Team
- Solo build
- Evidence
- 40 deterministic cases · complete traces
01 Problem
If the model grades the model, you will ship a confident wrong tool call. An ops agent on a warehouse API needs a judge that does not float.
02 Constraints
- 01The agent only acts through Clearbay MCP.
- 02Forty rule-scored cases. The model is never the judge.
- 03Tenant isolation, read-only writes, extra JSON fields, vague prompts must fail closed.
- 04Default router is deterministic. OpenAI tools are optional.
03 Architecture
04 Decisions
Eval harness is the product, not a later test.
Why Tenant leaks and extra fields are scoring rules, not vibes.
Tradeoff Cases are hand-written. That is cheaper than a wrong wave release.
Deterministic router by default.
Why You can demo isolation without an API key. The optional model is never the judge.
Tradeoff Vague language still needs a clarification case, not a guess.
Traces are first-class.
Why When a case fails, you read the tool call, not a chat bubble.
Tradeoff More logging than a toy agent. Necessary for an on-site loop.
05 Failure
Prompt tries another tenant.
Isolation case fails the run. Tool never fires cross-tenant.
Read role attempts a write.
Refusal case. MCP write tools require ROLE_OPS.
Extra JSON fields on a tool call.
Schema case. Unknown fields rejected.
Vague prompt.
Clarification case. No speculative write.
06 Object
07 Inspect
08 What I would improve next
- Generate adversarial and property-based cases, then use mutation testing to prove the rules catch regressions.
- Build a replay corpus across model and prompt versions instead of treating one pass rate as permanent.
- Add latency and cost budgets plus explicit human escalation for requests that cannot be scored safely.