Blog
Eight Rules for Benchmarking a Model You're About to Trust
Engineering, Sep 29, 2026, 12 min read

We ran one request through a new model five times. Median: 8.084 seconds.
The next day we ran the identical request again: same payload, same model, same code path. Median: 0.669 seconds.
Both numbers are real. Neither is wrong. And if we had run that request once, like most model comparisons do, we would have published whichever one we happened to get.
Swapping the model behind a decision looks like a small change. A prompt, a client, a parsed response, an afternoon. What makes it not small is where these decisions sit: route a request to a fast lookup when it needed a full investigation and the user gets a confident, shallow answer. Decide an agent can service a task when it holds credentials for one cluster out of nine, and the failure surfaces much later, somewhere that looks unrelated.
These are the eight rules we put between a decision model and production. We built them while evaluating TypeSafe's Jev against our own LLM classifiers, and every one is here because skipping it would have let us believe something comfortable.
1. Freeze real requests, not written ones
Every benchmark starts from real traffic. For the capability checker that was 10 real Dev EU requests spanning 2026-09-15 to 2026-09-23: 15 checks, 36 tasks, 215 agent calls. Not requests we wrote to be representative. Requests that happened, carrying the context they actually carried.
Then freeze them. The exact capability text, datasource metadata and task batch from each call, replayed byte for byte on every run of every arm. If one side sees slightly different context the comparison is void, and nothing in the output will tell you.
Ten requests is a small sample, and we say so wherever a result is close. It is small because every row needs a human to decide what the right answer was, and ten reviewed rows beat two hundred unreviewed ones. Sample size is the honest reason we decline to call a near-tie.
2. Run everything three to five times
Five repeats per request per arm on the capability checker, three on the token benchmarks, 195 total valid runs on intent routing.
Repeats are not about averaging out noise. They are about whether a model is consistent, which for a routing decision matters nearly as much as being right.
Identical inputs, five times. Our checker runs at temperature 1 and changed 5.3% of its decisions; Jev changed 0.8%. One run per request cannot tell "right 80% of the time" apart from "right 80% of the time by flipping," and those are different systems with different failure modes.
Then there is the case from the top of this post:
Unchanged payload, unchanged model. The first five runs sat at a median of 8.084 seconds with four of them between 7.379 and 14.864. The recheck the next day sat at 0.669, slowest 1.163.
We publish the combined median of all ten, 1.072 seconds, and we publish the split beside it, because the cause of that first batch is still unknown. Picking either half would have been a story. Publishing both is the measurement.
3. Record every run, and leave the gaps as gaps
Each run saves its decisions, the route taken, token counts, billed cost, wall duration, and the provider's response ID.
The response ID is the field everyone skips and the one that makes the rest trustworthy. It means any number in any summary can be walked back to the exact provider call that produced it, months later, by someone who was not in the room. Without it you have a spreadsheet that is true because you say so.
Missing measurements stay blank. Not zero, not an estimate. Our intent-routing benchmark has 195 valid records but only 135 saved timings, because 60 earlier runs predate us capturing duration. So the timing charts are drawn from 135 points and say so, rather than quietly filling 60 holes with something plausible.
4. Publish the runs you threw away
Two exclusions from intent routing, both written into the result:
Four attempts at one request ran against the wrong Gemini version. Excluded, re-run. One Jev call returned HTTP 520; we retried it successfully and excluded the failed attempt from that request's five valid repeats.
Neither is interesting, which is the point. Exclusions are suspicious only when they are invisible, and a benchmark reporting zero discarded runs is usually one that was not watched closely. Writing down the boring ones is what makes it credible when you say there were no others.
5. Say what the numbers do not prove
This is the section I would most like other people's benchmarks to have.
For the capability checker, both flows ran from inside the same pod through the same gateway, on the same frozen inputs, inside an eleven-minute window. That is a fair latency comparison and we present it as one.
For intent routing it is not, and we label it everywhere it appears. Our old-flow timer covers classifier initialisation, three parallel calls and replay cleanup. The Jev measurement covers an HTTP POST and excludes key lookup. Some runs were measured in-pod, some from a local client. Those are different quantities of work.
We also did not test final answer quality anywhere in the intent-routing benchmark. We measured which route each configuration chose. Whether the answer that came back down that route was any good is a question we did not ask, and a reader who assumes otherwise has been misled by us, not by themselves.
If a benchmark has no limitations section, it has limitations.
6. Agreement with production is not correctness
The subtlest trap here, and the one that nearly cost us a working model.
When you replay a real request you know what your current system decided. Treating that as truth is easy to compute and feels objective. It measures the wrong thing. Your current system's decisions are observed behavior, not reviewed answers. Ours had been handing GKE tasks to the AWS agent for months. Score agreement and the ceiling for any new model is reproducing your mistakes exactly, while every case where it is right and you were wrong counts against it.
So agreement is a pointer to where to look, never a score. For a score you need a key.
Ours took 4 read-only audits of each agent's real tools, deployed prompts, live datasource configuration and the benchmark traces. 134 (task, agent) pairs labelled CAN, PARTIAL, CANNOT or UNSURE, each citing its evidence. No LLM judged any label: a model grading a model's routing decision inherits exactly the ambiguity you are trying to measure.
Two things only an audited key can show you. Two of our tasks have no capable agent at all, so the correct answer is to assign nobody, something agreement scoring cannot even represent, because production assigned somebody. And PARTIAL had to exist: an agent that keeps two hours of events when the task wants seven days is neither a yes nor a no, and forcing it into a binary scores a reasonable refusal as an error.
Here is what trusting agreement would have cost us. Jev's first intent-routing configuration agreed with production on 27 of 50 runs. As a score, that is a rejection. As a pointer, it said go look at requests 4, 5, 7 and 10: all counts and lists, all answering from context instead of researching, in every single run. The cause was our own routing rule, which never said a complete list requires research. We rewrote the rule. Agreement went to 85 of 85.
We would have rejected a model that worked and kept the defect that caused the disagreement.
7. Trace every error to a cause
The rule that changed the conclusion of the whole evaluation.
Counting errors is not enough. Every wrong decision gets traced to why, and the categories have to include ones that do not blame the model, or you will never find out that the model was not the problem.
Three routers sat within five points of each other around 70% precision. Counted, that reads as "the new model is no better than what we have." Attributed, it reads completely differently: router error, where the description was clear, is 0.4 per run at worst and zero for three of the four flows. Everything else traces to capability descriptions claiming abilities the agents lacked, omitting which cluster they covered, or to a render step that leaked out of our planner.
We fixed the descriptions and re-ran. Our own incumbent checker gained 14.5 points of precision from a prompt-only change. Without the attribution step we would have shipped a post comparing two models and never found the actual bottleneck.
Attribution drags up things outside the benchmark, too. Classifying 507 decisions means reading code, and we found a formatter appending a colon to empty format specs (so every literal {…} in every prompt reached the model malformed) and a cache keyed on datasource UID but not instance, quietly sharing one catalog between two environments. Neither is a benchmark result. Neither would have been found by counting.
8. Report the curve, not a point
A scored model has no single accuracy. At threshold 0.4 it is 91% precise covering 91% of tasks; at 0.6, 97% precise covering 84%. Same model, same requests, same run. Only the cutoff moved.
Report one number and you have hidden a product decision inside it. Report the curve and the decision stays where it belongs: take coverage for a reversible suggestion a person reviews, take precision for a decision feeding something expensive and let it abstain.
One caveat we hold ourselves to: both of those thresholds were chosen on the same ten requests they are measured on. They are fitted until a fresh sample confirms them, and that belongs in the result, not a footnote.
The whole harness
- Real requests from traces, frozen with their context, replayed byte for byte.
- Three to five repeats per request per arm. Consistency is a result, not noise.
- Decisions, tokens, cost, duration and response ID per run. Gaps stay gaps.
- Exclusions published, including the boring ones.
- An explicit statement of what the numbers do not prove.
- Agreement as a pointer, never a score. Build an audited key, allow PARTIAL, let no model grade it.
- Every error traced to a cause, with categories that can exonerate the model.
- The precision-coverage curve, not a point, and a note when the thresholds were fitted.
None of this is sophisticated, which is deliberate: it is cheap enough to run for a decision as small as "which agent can service this task." Two of the three most valuable things it has told us were not about models at all. One was that sixteen text files were wrong. The other was that a rule we wrote ourselves was vague.
That is the argument for a harness. Not that it ranks vendors. That it tells you when you are about to buy a model to fix a problem a text file would have fixed.
The evaluation this came out of, with the full measurements and every caveat, is in We Benchmarked Jev Against Our Own Classifiers.