Mynd LabsResearchAll of Mynd Labs ↗
Research division
CollaboratePressContact

Research / Evaluation

Proof

Tests a field expert would actually sign off on.

Early researchEvaluationBenchmarksField experts

In brief

Benchmarks measure what is convenient and then saturate. Proof builds evaluations from real work, written with the people who do that work.

expertsign-off

Illustration of the idea. Not a result.

The problem

A number on a leaderboard says little about whether a system is safe to use in a hospital, a courtroom or a warehouse.

Benchmarks are easy to publish and hard to trust. They measure what is convenient, saturate quickly, and rarely resemble the work a person actually needs done.

Approach

How we are going about it.

  1. 01

    Start from real tasks

    Each evaluation begins with a task from a field, described by someone who does it, including what a bad answer costs.

  2. 02

    Expert sign-off as a criterion

    A result counts only if a practitioner would accept it. We record the criterion with the test.

  3. 03

    Retire what saturates

    Evaluations carry a life span. When a test stops separating systems it is retired and replaced.

Open questions

  • How many expert judgements are enough for a stable score?
  • How do you keep an evaluation from leaking into training data?
  • What does a fair comparison look like between a model and a system built around it?

What we aim to publish

  • Field-specific evaluation sets, built with practitioners
  • An open harness for running them
  • Methods notes explaining every choice

Limits. Expert judgement is slow and not perfectly consistent. Proof will report disagreement between experts instead of hiding it.

Landscape

What already exists, and where we start.

MMLU and HELM are broad and useful, and MMLU shows what saturation looks like. SWE-bench starts from real tasks, which is the direction we want. What we add is the practitioner's sign-off as a recorded criterion, and a planned end of life for each test.

  1. Paper
    Hendrycks et al. (2020). Measuring Massive Multitask Language Understanding.

    A widely used multiple-choice benchmark across 57 subjects. A useful case study in benchmark saturation.

  2. Paper
    Liang et al. (2022). Holistic Evaluation of Language Models.

    A broad evaluation framework from Stanford CRFM that reports many metrics across many scenarios.

  3. Paper
    Jimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    Evaluates models on real software issues instead of synthetic questions. Close to our idea of starting from real work.

  4. Paper
    Mitchell et al. (2019). Model Cards for Model Reporting.

    A documentation format for what a model is for and where it fails.

Work on Proof with us.

hello@myndlabs.tech