Hub
AI Benchmarks
Catalogue of evaluation suites tracked across the model graph. The benchmark itself is what we verify here; per-model scores are not yet verified by this platform and are deliberately not published.
MMLU-Pro
Verifiedknowledge
A harder, more reasoning-focused successor to MMLU covering broad academic and professional knowledge.
- Definition
- verified
- Per-model scores
- Data not yet verified.
GPQA Diamond
Verifiedreasoning
Graduate-level questions across physics, chemistry, and biology designed to be 'Google-proof'.
- Definition
- verified
- Per-model scores
- Data not yet verified.
SWE-bench Verified
Verifiedcoding
Real-world software engineering tasks sourced from GitHub issues, with verified solutions.
- Definition
- verified
- Per-model scores
- Data not yet verified.
AIME
Verifiedmath
American Invitational Mathematics Examination problems used to measure mathematical reasoning.
- Definition
- verified
- Per-model scores
- Data not yet verified.
HumanEval
Verifiedcoding
Classic functional code generation benchmark covering 164 hand-written Python problems.
- Definition
- verified
- Per-model scores
- Data not yet verified.