Research guide
Why benchmark scores need source discipline
Benchmark scores are useful only when their provenance, dataset, and prompt protocol are documented. This page explains why WebmasterID Models does not republish unsourced provider-reported scores.
Last updated: 2026-05-21
Benchmark vs benchmark score
A benchmark is an evaluation specification — a dataset, a prompt protocol, a scoring rule. A benchmark score is a single model's result against that specification. The two are different objects in the data model. The catalogue records the structural benchmark — its name, category, description, and stable slug — but does not republish per-model scores without source-level verification.
The category enum is fixed: reasoning, coding, math, knowledge, multimodal, agentic. New benchmark types are added by extending the enum and the BenchmarkEntity record, not by stuffing data into a free-text notes field.
Provider-reported vs independent
A provider quoting their own model's benchmark score is a vendor-reported claim, on the same epistemic footing as a vendor-reported status indicator. It can be useful as a colour signal but it is not an independent measurement. The catalogue treats provider-reported scores the same way it treats vendor-reported status: surface only with explicit attribution, never as a property of the model itself.
Independent leaderboard scores (Chatbot Arena, MMLU-Pro, SWE-bench Verified runs hosted by third parties, etc.) are legitimate primary sources when they document the evaluation protocol, the model snapshot evaluated, and the date. Where such a citation exists, a score can be encoded on the ModelEntity as a verified field. Until it exists, no score appears.
Dataset contamination
Public benchmark datasets are reachable from public training data. Once a model is trained on enough of the open web, its score on that benchmark is no longer a clean measurement of generalisation — it's a measurement of memorisation plus generalisation, in some unknown mix. Published numbers on widely-known benchmarks tend to drift upward over time across all vendors for this reason. Without an explicit decontamination protocol, a single high score is hard to interpret.
We do not have a clean way to encode this as a per-score field today, which is one of the reasons the published-score surface stays narrow.
Prompt and version drift
A benchmark score depends on the prompt protocol used at evaluation time. Two runs of the same benchmark with different system prompts, different temperature settings, or different chain-of-thought conventions can produce materially different scores on the same model. Provider reports rarely document the exact protocol used; community leaderboards usually do. The catalogue treats both as sources but applies the same provenance rules.
Model versions also drift. A rolling alias like claude-opus-4-7 can resolve to different snapshots over time; a benchmark score against a rolling alias is only meaningful when paired with the evaluation date. Pinned snapshot IDs (e.g. claude-opus-4-20250514) are stable; aliases are not.
Lifecycle changes between scores and publication
A score published the week a model went GA can become stale quickly. Quantisation changes, runtime upgrades, and retirement announcements all shift what a leaderboard number means in production. The catalogue records lifecycle status and retirement dates as verified fields — see /research/model-selection — so a published score can be read against its model's current state instead of in isolation.
Limitation matrix — why benchmark scores warrant care
| Visible symptom | Effect on interpretation | |
|---|---|---|
| Dataset contaminationPublic benchmark in public training data | Score drifts upward over time across vendors. | Memorisation + generalisation are mixed; the number is no longer a clean signal of capability. |
| Prompt varianceDifferent system prompts / temperature / CoT conventions | Same model, same dataset, different scores depending on protocol. | Single-number reports without protocol are not reproducible. |
| Model version driftRolling aliases vs pinned snapshots | Score from an alias on day X reflects whichever snapshot served then. | Without snapshot date, scores cannot be replicated. |
| Provider-reported vs independentWho ran the eval matters | Provider blog post quoting their own score. | Same epistemic footing as a vendor-reported status indicator: useful as colour, not a property of the model. |
| Task / domain mismatchBenchmark domain ≠ production workload | High score on MMLU; low real-world performance on internal docs. | Score signals nothing about workload-specific capability. |
| Reproducibility gapClosed eval harness, undocumented seed | Two runs yield different scores; no published harness. | Cannot be re-verified independently. |
What the catalogue publishes today
The structural benchmark list at /benchmarks records benchmark names, categories, and descriptions. No per-model scores are published. When scores eventually appear, each will carry: a primary-source citation, the evaluation date, the model snapshot evaluated, the dataset version, and a free-text protocol note. The integrity guard suite already refuses to ship JSON-LD with unverified benchmark fields; that constraint extends to the page body when scores are added.
What this page assumes is verified
Verified today
Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.
Benchmark categories tracked
Reasoning, coding, math, knowledge, multimodal, agentic — declared in the BenchmarkEntity schema with stable slugs. See /benchmarks for the live list.
Zero model benchmark scores published
No ModelEntity in the catalogue carries a verified benchmark score field today. Provider-reported scores are not republished without independent verification; independent leaderboard scores require a primary-source citation that is not yet on record.
Honest gaps
Data gaps
Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.
Independent benchmark harness
The platform does not run its own benchmark suite. Scores would require either a verified independent leaderboard citation or a documented in-house evaluation protocol — neither is in scope today.
Benchmark dataset metadata
When a score is eventually published, it will need to carry the dataset version, prompt protocol, and evaluation date alongside the number. None of these fields exist yet on the data layer.
Continue