Skip to content
WebmasterID

Learn · model fundamentals

How to choose an AI model

A workflow for picking which AI model to test next — start from your use case, inspect verified fields, then export an evidence brief. Never start from a leaderboard.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

Start from the use case, not the leaderboard

Most AI model selection failures come from starting at a leaderboard and working backwards. A model that tops a generic benchmark may still fail your workload because of context window limits, modality gaps, lifecycle status, regional availability, or pricing semantics that look fine in marketing copy but matter in production.

The catalogue is structured to support the inverse workflow: define the use case first, then narrow the shortlist to models with verified fields that match, then compare them side by side, then export the evidence so the rest of the team can review.

Why this matters

Model decisions are infrastructure decisions. They shape your cost structure, your latency budget, your data-handling story, and your compliance posture. Treating that decision as "whichever model the blog post said is best" gives you no audit trail when the team asks why you chose it.

What to verify before integrating any model

  • The model's lifecycle status — is it active, preview, deprecated, or retired? Retirement dates often appear months before they bite.
  • The model's context window — does it cover your full prompt plus retrieved context plus expected output?
  • The model's max output tokens — separate from context window. A 1M-token context with 8k output limits some workloads.
  • The model's input modalities — text only, or images, audio, video too?
  • The pricing reference — is it first-party or hosted? Different sources, different freshness, different terms.
  • The verification status of every field above — does the catalogue have a primary-source citation on record?
  • The data gaps the catalogue marks with the unverified-data label — those are the questions you still need to answer externally.

Example: verified context windows in the catalogue

Below is a small slice of the catalogue showing how verified fields render. The table is for inspection — the catalogue does not assert that any of these models is better for your workload than the others.

Verified examples · Context window

Verified context windows pulled directly from the typed data layer. Click a row to read the full record and its citations.

ModelProviderContext windowSource state
Claude Opus 4.7Anthropic1,000,000 tokensverified · citation on record
Claude Sonnet 4.6Anthropic1,000,000 tokensverified · citation on record
Gemini 2.5 ProGoogle1,048,576 tokensverified · citation on record
DeepSeek V4 ProDeepSeek1,000,000 tokensverified · citation on record
Mistral Large 3Mistral256,000 tokensverified · citation on record

Inspection only. The catalogue does not rank these models on the field above.

Common mistakes

  • Treating context window as a hard fit indicator alone.

    A model with a huge context window can still degrade on long inputs. Context size is a necessary condition, not a sufficient one.

  • Reading hosted pricing as the model creator's pricing.

    Hosted pricing is set by the hosting platform. The same model can have very different pricing semantics across creators and hosts.

  • Ignoring lifecycle status.

    Integrating a deprecated snapshot weeks before its retirement date locks the team into a migration immediately after launch.

  • Treating a missing value as zero.

    An unverified field is not 'no support' — it's 'no primary-source citation on record yet'. Confirm externally before assuming either way.

Apply this workflow

The catalogue exposes the steps as workspaces. Each link below carries the use case forward so you only fill the form once.

Apply this workflow

Practise this lesson

These exercises route the lesson concept through the verified-data product surfaces. Each one ends with a concrete artifact you can share.

Data gaps to watch

Some catalogue values intentionally remain unverified — most commonly latency, throughput, uptime, and any provider that does not publish machine-readable docs. These gaps are surfaced through a single canonical unverified-data label rather than substituted with estimates. Treat that label as a prompt to verify externally, not as a signal that the model lacks the capability.

Related pages

Sources and freshness

Every value in the example table above is wrapped with a primary-source citation and a retrieval date. Open any model page to read the underlying URLs, or visit /sources for the citation registry. Citations age; the reverification queue lists what is due for re-check.

Teaching example

Illustrative — not a recommendation.

Situation: A team needs to add an AI model to summarise customer support tickets weekly. They have not yet defined the model selection workflow.

Decision to make: Which 2–3 candidate models should we evaluate this quarter, and what evidence will we record?

Verified fields that matter:

  • Lifecycle status (active only)
  • Verified context window vs typical ticket length
  • Pricing reference + retrieval date
  • Hosted vs first-party availability
  • Source freshness on the citations

What to inspect next:

Weak vs better approach

Weak approach

  • Open the most recent vendor blog post.
  • Pick the model the blog headlines.
  • Skip lifecycle / pricing / source verification.
  • Ship and hope the snapshot does not rotate during launch.

Better approach

  • Name the use case and its acceptance rubric.
  • Filter the catalogue to active lifecycle + verified pricing.
  • Compare 2–4 candidates on verified fields side by side.
  • Export a Markdown brief listing fields, gaps, sources, freshness.

Why better: The weak approach optimises for novelty. The better approach optimises for an auditable trail your reviewer can read independently.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Decision brief excerpt — shortlist note

## Shortlist (illustrative)
Use case: support-ticket summarisation
Candidates: <model-A>, <model-B>, <model-C> (slugs from /select)

Verified fields per candidate:
- Lifecycle: active (cite source)
- Context window: <tokens> (cite source, retrievedAt)
- Hosted availability: yes / no
- Pricing reference: <unit> + retrievedAt

Open data gaps:
- Max output tokens not stated for <model-C> (unverified-data label)

Next step: open /compare/build for the three slugs and export brief.

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Frame the decision workflow before opening the catalogue.

  2. Step 2

    Apply in /select →

    Narrow a source-backed shortlist by use case + lifecycle.

  3. Step 3

    Verify in /sources →

    Read the primary-source citation for each verified field.

  4. Step 4

    Test with /lab →

    Run the prompt-testing playbook against the shortlist.

Review before moving on

  • I can name the use case and the acceptance rubric.
  • Every candidate's lifecycle field is verified active.
  • Every candidate has at least one verified pricing reference or an explicit gap noted.
  • I have a /select URL I can share with a teammate.
  • I have NOT named a winner — only candidates with evidence.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Picking which AI model is 'best' — that depends on your workload, not the catalogue.
  • Ranking models by price, latency, throughput, or uptime — the catalogue records observations, not rankings.
  • Certifying compliance — verification means a citation backed the value on the date recorded, not that the model meets any specific regulation.