Skip to content
WebmasterID

Hub

AI Benchmarks

Catalogue of evaluation suites tracked across the model graph. The benchmark itself is what we verify here; per-model scores are not yet verified by this platform and are deliberately not published.

  • MMLU-Pro

    Verified

    knowledge

    A harder, more reasoning-focused successor to MMLU covering broad academic and professional knowledge.

    Definition
    verified
    Per-model scores
    Data not yet verified.
  • GPQA Diamond

    Verified

    reasoning

    Graduate-level questions across physics, chemistry, and biology designed to be 'Google-proof'.

    Definition
    verified
    Per-model scores
    Data not yet verified.
  • SWE-bench Verified

    Verified

    coding

    Real-world software engineering tasks sourced from GitHub issues, with verified solutions.

    Definition
    verified
    Per-model scores
    Data not yet verified.
  • AIME

    Verified

    math

    American Invitational Mathematics Examination problems used to measure mathematical reasoning.

    Definition
    verified
    Per-model scores
    Data not yet verified.
  • HumanEval

    Verified

    coding

    Classic functional code generation benchmark covering 164 hand-written Python problems.

    Definition
    verified
    Per-model scores
    Data not yet verified.