Research note · 2026-10-06
Benchmarks saturate. Expert sign-off does not.
A benchmark is useful for a while. Then systems learn the format, scores bunch near the top, and the number stops telling you anything about the work.
The usual response is a harder benchmark. We think the better response is to change what counts. A result should count when a practitioner would accept it for real work, and the evaluation should record that criterion next to the test.
That means building evaluations with people who do the work, including what a bad answer costs them. A wrong contract clause and a wrong product description are not the same failure.
Expert judgement is slower and less consistent than an automatic score. We would rather report disagreement between experts than hide it behind an average.
Evaluations should also have a life span. When a test stops separating systems it should be retired and replaced, and that decision should be written down.
Related project: Proof