Skip to content
WebmasterID

Learn · comparison methodology

Why benchmark scores can mislead

Contamination, prompt variance, version drift — and why the catalogue does not publish provider-reported benchmark scores casually.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

Benchmark score vs benchmark definition

A benchmark name (think of the well-known reasoning, coding, math, knowledge, and instruction-following benchmarks) is a definition: a dataset, a scoring rubric, and an evaluation protocol. A benchmark score is one model's performance on that definition, measured by a specific party at a specific time on a specific snapshot. The two are not the same.

The catalogue records benchmark definitions as entities — name, category, description, citations to the canonical paper or repository. It does not publish per-model scores, because score quality depends entirely on the conditions under which the score was produced.

Why provider-reported scores are not published casually

Provider-published benchmark numbers are often impossible to reproduce independently. They are usually generated by the provider's own evaluation harness, on a specific model snapshot, with proprietary prompt templates. The catalogue's verification rules require either an independent primary source or sufficient methodology that the reader can reproduce the number themselves. Most provider-reported scores do not meet that bar.

Contamination, prompt variance, version drift

  • Contamination — the benchmark data may appear in the model's pretraining set. The model is then recalling, not reasoning. Most popular benchmarks have some contamination footprint by now.
  • Prompt variance — the same model answers the same benchmark question differently under different system prompts, sampling temperatures, or formatting. Two providers can publish "the same" benchmark with different numbers because they ran it differently.
  • Version drift — providers update their underlying weights without changing the model name. A benchmark run last quarter is not necessarily a benchmark run today.

What to look at instead of a published score

  • The benchmark's own definition page — what is it actually measuring?
  • Whether the model's documentation cites the benchmark with a methodology you can reproduce.
  • Whether the catalogue carries a citation to the methodology, separate from any score.
  • Whether your own workload's prompts and rubric exist — if not, you have no benchmark, regardless of which one the provider names.
  • Whether the model snapshot you are integrating matches the snapshot the score was produced on.

Common mistakes

  • Picking a model because its provider posted a higher score.

    The provider's score is a marketing artifact unless the methodology is independently reproducible.

  • Comparing scores across providers without aligning protocols.

    Different prompt templates, sampling parameters, and snapshots can change scores by 10+ points.

  • Treating contamination as a 'newer model' problem.

    Contamination has been documented across many well-known benchmarks for years. Newer is not always cleaner.

  • Reading version drift as 'an improvement'.

    Drift can also degrade performance on workloads the new snapshot was not optimized for.

Apply this workflow

Apply this workflow

Data gaps to watch

The absence of a benchmark score row in the catalogue is intentional. It means the catalogue's verification rules did not find an independently reproducible source, not that the model performs poorly. Treat the absence as "run your own evaluation", not "missing data".

Related pages

Sources and freshness

Benchmark definitions are stable; the canonical papers and repositories do not change often. The reverification queue only flags benchmark entities when a benchmark moves to a new canonical location (e.g. a paper revision, a repository archive).

Teaching example

Illustrative — not a recommendation.

Situation: A vendor blog cites a benchmark figure that appears stronger than the team's prior workload-specific evaluation. The team is asked whether to switch candidates based on the new figure.

Decision to make: Does the vendor's benchmark methodology meet our reproducibility bar, and how should the figure be treated in our brief?

Verified fields that matter:

  • Benchmark definition (catalogue entity)
  • Provider's published methodology
  • Snapshot tested
  • Our own workload acceptance rubric

Weak vs better approach

Weak approach

  • Switch candidates based on a vendor-cited benchmark figure.
  • Treat one provider's numbers as comparable to another's.
  • Skip the methodology check.
  • Assume newer snapshot = better benchmark.

Better approach

  • Confirm the benchmark definition the vendor cited matches the one the team cares about.
  • Check whether the methodology is independently reproducible.
  • Recognise prompt-variance and version-drift effects in your interpretation.
  • Rely on workload-specific tests; benchmark figures are not selection evidence on their own.

Why better: Without independently reproducible methodology, a benchmark figure is a marketing artifact. The better approach uses the benchmark definition as a vocabulary, not as a comparable score.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Benchmark interpretation note

## Benchmark interpretation (illustrative)
Cited figure: <verbatim quote>
Source: <vendor URL>
Benchmark definition referenced: <name + canonical link>

Methodology check:
- Snapshot tested: <slug>
- Prompt templates published: <yes / no>
- Independent reproduction available: <yes / no>

Interpretation: <treat as marketing artifact / treat as reproducible signal>
Decision impact: <none, until workload-specific tests support it>

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Understand why catalogue benchmarks are definitions, not scores.

  2. Step 2

    Apply in /benchmarks →

    Read the benchmark definition before reading any cited figure.

  3. Step 3

    Verify in /sources →

    Trace any cited figure back to its primary source + methodology.

  4. Step 4

    Test in /lab →

    Run your own evaluation prompt set instead of trusting a published score.

Review before moving on

  • I can name the benchmark definition referenced — not just the figure.
  • I checked whether the methodology is independently reproducible.
  • I have NOT used the figure as selection evidence on its own.
  • Workload-specific tests will drive the decision.
  • Any figure in the brief is paired with the methodology caveat.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Publishing benchmark scores. The catalogue does not, and the lesson does not.
  • Asserting which model is best on any benchmark.
  • Replacing your own workload-specific evaluation work.