In brief
Benchmarks measure what is convenient and then saturate. Proof builds evaluations from real work, written with the people who do that work.
Illustration of the idea. Not a result.
The problem
A number on a leaderboard says little about whether a system is safe to use in a hospital, a courtroom or a warehouse.
Benchmarks are easy to publish and hard to trust. They measure what is convenient, saturate quickly, and rarely resemble the work a person actually needs done.
Approach
How we are going about it.
- 01
Start from real tasks
Each evaluation begins with a task from a field, described by someone who does it, including what a bad answer costs.
- 02
Expert sign-off as a criterion
A result counts only if a practitioner would accept it. We record the criterion with the test.
- 03
Retire what saturates
Evaluations carry a life span. When a test stops separating systems it is retired and replaced.
Open questions
- How many expert judgements are enough for a stable score?
- How do you keep an evaluation from leaking into training data?
- What does a fair comparison look like between a model and a system built around it?
What we aim to publish
- Field-specific evaluation sets, built with practitioners
- An open harness for running them
- Methods notes explaining every choice
Limits. Expert judgement is slow and not perfectly consistent. Proof will report disagreement between experts instead of hiding it.
Landscape
What already exists, and where we start.
MMLU and HELM are broad and useful, and MMLU shows what saturation looks like. SWE-bench starts from real tasks, which is the direction we want. What we add is the practitioner's sign-off as a recorded criterion, and a planned end of life for each test.
- PaperHendrycks et al. (2020). Measuring Massive Multitask Language Understanding.
A widely used multiple-choice benchmark across 57 subjects. A useful case study in benchmark saturation.
- PaperLiang et al. (2022). Holistic Evaluation of Language Models.
A broad evaluation framework from Stanford CRFM that reports many metrics across many scenarios.
- PaperJimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Evaluates models on real software issues instead of synthetic questions. Close to our idea of starting from real work.
- PaperMitchell et al. (2019). Model Cards for Model Reporting.
A documentation format for what a model is for and where it fails.
Notes