Skip to content
Hillary Uzomba
Verdict August 20, 2026

The benchmark that lied: measuring a cache, not a model

A repeated prompt came back 182 times faster. Nothing about the model had changed. Here is how to spot the error and how to measure properly.

2 min read Hill

Photo by Arnold Francisca on Unsplash

A benchmark is only ever a claim about one machine on one afternoon. Run the same prompt twice and the second one comes back faster, and if you are not careful you will write that number down and call it a result.

This is the trap I keep finding in my own notes. It is worth setting out plainly, because the error is not small. In one round of testing on a single machine, a repeated prompt appeared to run 182 times faster than its first pass. Nothing about the model had changed. What I had measured was a cache.

What is actually happening

When a local runtime processes a prompt, it does two quite different things. First it reads the prompt itself, turning every token into the internal state the model needs. Then it generates, one token at a time, using that state.

The first part is expensive and highly parallel. The second is cheaper per token but stubbornly sequential. Most runtimes keep the result of the first part around, keyed on the prompt’s opening tokens, so that a conversation does not pay to re-read its own history on every turn. That is a sensible optimisation and it is why chat feels responsive.

It also means that the second time you send an identical prompt, the expensive half is skipped entirely. You are no longer timing the model. You are timing a lookup.

How to tell it is happening to you

The symptom is a speed-up too large to be real. Hardware improvements arrive in increments of tens of percent. If a change appears to have made something an order of magnitude faster, and you did not change the hardware, the quantisation or the context length, you have almost certainly measured your own cache.

The other tell is that the gain sits entirely in time-to-first-token while the generation rate barely moves. That is the shape of a skipped prefill, not a faster model.

Measuring it properly

  • Vary the prompt every run. Not the ending, the beginning. A cache keyed on a prefix is unbothered by a changed last sentence.
  • Report the two numbers separately. Prefill and generation answer different questions; averaging them hides both.
  • Discard the first run, then discard the idea of a single number. Report a range across several cold runs.
  • Say what was cached. A result without that note cannot be reproduced by anyone, including you in six months.

Why it matters more locally

Running models on your own hardware is mostly an argument about honesty: you can see the whole system, so you can say what you actually observed. That advantage is thrown away the moment a cached number goes into a comparison table. Worse, the error flatters whichever setup you tested last, which is usually the one you were hoping would win.

The fix costs nothing. Change the prompt, run it cold, write down both numbers. A slower honest figure is worth more than a fast one you cannot defend.