Research guide
Choosing an AI model: a verified-data approach
A practical, source-aware framework for selecting an AI model — covering provider, pricing, context window, output limits, modality, lifecycle, and reliability signals.
Last updated: 2026-05-21
Selection is not only benchmark score
Choosing an AI model for a production workload is a decision across at least eight independent axes. A leaderboard score — even one with strong provenance — can capture only a slice of the surface a real system depends on. WebmasterID Models is deliberately structured around the verifiable axes a builder can actually act on: provider attribution, pricing, context window, output limit, modality, lifecycle status, API surface, and observable reliability signals.
The page you are reading is not a buying recommendation. It is a checklist of the criteria that have stable, sourceable values across providers, together with pointers to the live data — so the same decision becomes auditable instead of opinion-driven.
The verifiable criteria
The following dimensions are first-class fields in the model catalogue. Every value is either verified against a primary source (with the citation visible on the model page) or rendered as the canonical unverified-data label. Nothing is interpolated.
- Provider. Who trains and serves the model. See /providers for the catalogue and /docs/provider-coverage for the dimensions we track per provider.
- Canonical API identifier. The model string you pass on the wire, plus any aliases or platform-specific identifiers (e.g. Bedrock / Vertex). Pinned snapshot IDs are recorded separately from rolling aliases.
- Pricing. Per-token input, output, cache write, cache read, batch-tier rates as the vendor publishes them. Each row carries its own unit semantics — see /research/api-pricing-methodology.
- Context window. Input token limit per request as published. Million-token claims warrant care — see /research/model-context-windows.
- Max output tokens. A separate dimension from context window; structured-output workloads and agentic loops are constrained here.
- Modality. Which input/output channels are documented (text, image, audio, video — separately for input and output).
- Lifecycle. Active, preview, deprecated, or retired. Deprecation announcements carry retirement dates and migration targets where the vendor publishes them.
- API surface. Endpoint shape, model parameter position, authentication conventions. Vendor docs change shape over time; WebmasterID Models records the snapshot used and links back to the source.
Lifecycle and deprecation
A model with a strong leaderboard score that is scheduled for retirement next quarter is, for production purposes, a different decision than one with the same characteristics on an active lifecycle. The catalogue surfaces lifecycle status and retirement dates as verified fields directly on each model page. The Claude Opus 4 (2025-05-14) record is the gold-standard example: it carries an explicitdeprecated status, retirement date, and migration target sourced from Anthropic's own deprecation table.
Historical entries (DeepSeek R1, Mistral Large 2) are kept as structural records with retired lifecycle so search engines and AI crawlers cannot accidentally treat them as live recommendations.
Reliability signals
The platform separates three kinds of reliability data:
- Vendor-reported status. The provider's own status feed, observed at a fixed cadence and clearly labelled as vendor-reported. See /status for the per-provider observer matrix and /research/ai-provider-status-monitoring for the policy.
- Independent HTTP probes. A request issued by WebmasterID against a public, non-inference endpoint (host root). A successful probe is a reachability signal, not an availability measurement.
- Computed uptime window. Not published yet. Requires durable observations over a meaningful window (currently a minimum of 24 stored samples). When it appears, it will be labelled precisely — "vendor-reported operational-sample rate" — never "uptime" without qualification.
A model selection that depends on availability should weight a long-running observation history against a fresh leaderboard score, but neither signal is sufficient alone.
How unknown data is treated
The platform never interpolates, estimates, or averages missing values. When a field is not yet verified against a primary source, it renders the canonical unverified-data label and the model row reports a partial verification status. This is deliberately conservative: a fabricated number is worse than a clearly missing one for selection workflows that want to be auditable.
For example, Mistral Large 3 has a verified API string and lifecycle, but its per-model spec card returns 404 to automated retrieval, so its context window and pricing remain null. A reader considering Mistral can see exactly which fields are open and can either (a) confirm the missing values in a manual browser pass or (b) deprioritise the model until the gap closes.
Selection checklist
| Field on /models | Where to confirm | |
|---|---|---|
| Provider attribution | providerSlug | /providers |
| Canonical API ID + aliases | apiIdentifiers.canonical / alias | Model detail page → API identifiers section |
| Lifecycle (active / deprecated / retired) | lifecycle.status | /models?lifecycle=active |
| Context window | contextWindow | /research/model-context-windows |
| Max output tokens | maxOutputTokens | /research/model-output-limits |
| Modality | modality (text-in, image-in, audio-in, video-in, text-out, …) | Model detail page → Modality field |
| Pricing units (input/output/cache/batch) | pricing tiers + units | /research/api-pricing-methodology |
| Status observations | Provider observers + sample threshold | /research/ai-provider-status-monitoring |
| Source coverage | Per-model citations + per-provider coverage | /coverage |
What not to optimise for
- Headline benchmark score on its own. Without protocol / dataset / snapshot date, a score does not predict your workload — see benchmark limitations.
- Unverified latency claims. The catalogue does not publish a request-latency number for any model; the only latency-shaped value recorded is the status probe's fetch wall-clock, which is not API latency.
- Unsourced cost estimates. Cost depends on tokenizer + cache-hit rate + workload mix; published per-token rates are inputs, not totals.
Next steps
A shortlist workflow that uses this catalogue typically starts at /models (use the provider and lifecycle filters), narrows by modality and context window, and confirms pricing on /pricing. Two-sided verified comparisons live under /compare, each linked back to the source trail. Coverage state for every provider is on /coverage, and the underlying citations are indexed at /sources.
What this page assumes is verified
Verified today
Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.
Three providers have end-to-end model verification
Anthropic (Claude Opus 4.7, Sonnet 4.6, Haiku 4.5), Google (Gemini 2.5 Pro), and DeepSeek (V4 Pro). See /coverage for the per-provider matrix and /sources for the underlying citations.
Mistral is partially verified
Mistral Large 3 has API string and lifecycle verified; per-model spec card pages 404 to automated retrieval, so context window / output limit / modality / pricing remain unverified.
OpenAI is blocked on automated retrieval
platform.openai.com returns HTTP 403 to non-interactive clients. The GPT-5 catalogue row is a structural entry only; no metric has been published.
Honest gaps
Data gaps
Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.
No quality ranking
WebmasterID Models does not publish 'best model' claims or composite quality scores. Quality is task-dependent and not captured by any single number we are willing to source.
No independent benchmark scores
Provider-reported scores live with the provider; independent leaderboard scores are not republished unless a primary source is recorded. See /research/benchmark-limitations.
No latency or uptime ranking
Vendor-reported status is monitored for two providers (see /status) but no uptime percentage is published. Request-latency comparisons require a measurement harness we have not yet wired.
Continue
Related pages
API pricing methodology
How to read the rows on /pricing and why provider pricing cannot always be normalised.
Context windows in practice
Why a million-token context window on one provider is not the same as a million on another.
Model page schema
Reference for every field on a ModelEntity record.
Coverage
Per-provider verification matrix and retrieval audit log.