Skip to content
WebmasterID

Research guide

Why benchmark scores need source discipline

Benchmark scores are useful only when their provenance, dataset, and prompt protocol are documented. This page explains why WebmasterID Models does not republish unsourced provider-reported scores.

Last updated: 2026-05-21

Benchmark vs benchmark score

A benchmark is an evaluation specification — a dataset, a prompt protocol, a scoring rule. A benchmark score is a single model's result against that specification. The two are different objects in the data model. The catalogue records the structural benchmark — its name, category, description, and stable slug — but does not republish per-model scores without source-level verification.

The category enum is fixed: reasoning, coding, math, knowledge, multimodal, agentic. New benchmark types are added by extending the enum and the BenchmarkEntity record, not by stuffing data into a free-text notes field.

Provider-reported vs independent

A provider quoting their own model's benchmark score is a vendor-reported claim, on the same epistemic footing as a vendor-reported status indicator. It can be useful as a colour signal but it is not an independent measurement. The catalogue treats provider-reported scores the same way it treats vendor-reported status: surface only with explicit attribution, never as a property of the model itself.

Independent leaderboard scores (Chatbot Arena, MMLU-Pro, SWE-bench Verified runs hosted by third parties, etc.) are legitimate primary sources when they document the evaluation protocol, the model snapshot evaluated, and the date. Where such a citation exists, a score can be encoded on the ModelEntity as a verified field. Until it exists, no score appears.

Dataset contamination

Public benchmark datasets are reachable from public training data. Once a model is trained on enough of the open web, its score on that benchmark is no longer a clean measurement of generalisation — it's a measurement of memorisation plus generalisation, in some unknown mix. Published numbers on widely-known benchmarks tend to drift upward over time across all vendors for this reason. Without an explicit decontamination protocol, a single high score is hard to interpret.

We do not have a clean way to encode this as a per-score field today, which is one of the reasons the published-score surface stays narrow.

Prompt and version drift

A benchmark score depends on the prompt protocol used at evaluation time. Two runs of the same benchmark with different system prompts, different temperature settings, or different chain-of-thought conventions can produce materially different scores on the same model. Provider reports rarely document the exact protocol used; community leaderboards usually do. The catalogue treats both as sources but applies the same provenance rules.

Model versions also drift. A rolling alias like claude-opus-4-7 can resolve to different snapshots over time; a benchmark score against a rolling alias is only meaningful when paired with the evaluation date. Pinned snapshot IDs (e.g. claude-opus-4-20250514) are stable; aliases are not.

Lifecycle changes between scores and publication

A score published the week a model went GA can become stale quickly. Quantisation changes, runtime upgrades, and retirement announcements all shift what a leaderboard number means in production. The catalogue records lifecycle status and retirement dates as verified fields — see /research/model-selection — so a published score can be read against its model's current state instead of in isolation.

Limitation matrix — why benchmark scores warrant care

Common limitation patterns affecting benchmark scores
Visible symptomEffect on interpretation
Dataset contaminationPublic benchmark in public training dataScore drifts upward over time across vendors.Memorisation + generalisation are mixed; the number is no longer a clean signal of capability.
Prompt varianceDifferent system prompts / temperature / CoT conventionsSame model, same dataset, different scores depending on protocol.Single-number reports without protocol are not reproducible.
Model version driftRolling aliases vs pinned snapshotsScore from an alias on day X reflects whichever snapshot served then.Without snapshot date, scores cannot be replicated.
Provider-reported vs independentWho ran the eval mattersProvider blog post quoting their own score.Same epistemic footing as a vendor-reported status indicator: useful as colour, not a property of the model.
Task / domain mismatchBenchmark domain ≠ production workloadHigh score on MMLU; low real-world performance on internal docs.Score signals nothing about workload-specific capability.
Reproducibility gapClosed eval harness, undocumented seedTwo runs yield different scores; no published harness.Cannot be re-verified independently.

What the catalogue publishes today

The structural benchmark list at /benchmarks records benchmark names, categories, and descriptions. No per-model scores are published. When scores eventually appear, each will carry: a primary-source citation, the evaluation date, the model snapshot evaluated, the dataset version, and a free-text protocol note. The integrity guard suite already refuses to ship JSON-LD with unverified benchmark fields; that constraint extends to the page body when scores are added.

What this page assumes is verified

Verified today

Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.

  • Benchmark categories tracked

    Reasoning, coding, math, knowledge, multimodal, agentic — declared in the BenchmarkEntity schema with stable slugs. See /benchmarks for the live list.

  • Zero model benchmark scores published

    No ModelEntity in the catalogue carries a verified benchmark score field today. Provider-reported scores are not republished without independent verification; independent leaderboard scores require a primary-source citation that is not yet on record.

Honest gaps

Data gaps

Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.

  • Independent benchmark harness

    The platform does not run its own benchmark suite. Scores would require either a verified independent leaderboard citation or a documented in-house evaluation protocol — neither is in scope today.

  • Benchmark dataset metadata

    When a score is eventually published, it will need to carry the dataset version, prompt protocol, and evaluation date alongside the number. None of these fields exist yet on the data layer.

Continue

Related pages