Goal
Render a side-by-side comparison page where the only verified fields you look at are context window and max output tokens — and articulate what those values do not tell you.
Prerequisites
- Read /learn/context-window so the difference between context window and max output is clear.
- Have a list of 3–4 candidate model slugs in mind (your shortlist from the previous exercise, or any active models from /models).
Step-by-step exercise
Open the comparison builder
Visit /compare/build and select 3 or 4 candidate models. Optionally set the use-case filter.
Expected outcome: The page renders columns for your selected models. The URL captures the selection.
Read context window and max output rows
Locate the verified context window and max output token rows. Note which columns render a verified value vs the unverified-data label.
Expected outcome: You can describe the spread of context windows and max output limits across the columns — without naming a winner.
Open the citation for the largest context window
Click through to the model page for the row with the largest verified context window. Read the citation it points at.
Expected outcome: You can name the primary source the catalogue cited for that value and when it was retrieved.
Note what context size does not guarantee
Write down (or paste into your notes) the workload-specific behaviours the catalogue does not measure: deep-recall accuracy, instruction-following at length, prompt-size pricing tier impact.
Expected outcome: Your notes capture at least three things you still need to test in your own environment.
Completion checklist
Completion checklist
- The comparison URL is saved.
- You can describe the spread of context windows.
- You read the source URL for at least one verified value.
- You did NOT pick the model with the largest context window as the winner.
Evidence artifact
At the end of this exercise you should have: A /compare/build URL with 3–4 model slugs that any teammate can open to see the same comparison.
Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.
Related workflow routes
- /select — narrow the source-backed shortlist.
- /compare/build — render verified fields side by side.
- /briefs/build — generate the evidence decision brief.
- /sources — every primary-source citation, by provider.
- /coverage — per-provider verified-field coverage.
- /reverification — sources due for manual re-check.
Exercise does not recommend a model — external testing still required.
Common mistake
Treating the candidate with the largest verified context window as the answer. Larger context does not equal better recall or better fit.
Example artifact
Illustrative example — not a recommendation. Substitute your own values when you run the workflow.
Example artifact
## Context comparison URL (illustrative)
/compare/build?models=<slug-A>,<slug-B>,<slug-C>&useCase=long-context-analysis
Spread observed:
- Context window: <range>
- Max output: <range>
Notes to capture before testing:
- Recall behaviour is workload-specific.
- Pricing tiers may scale non-linearly with prompt size.Substitute your real shortlist, slugs, dates, and values when you walk the exercise.
Repeat this exercise when
- A new candidate is added to the shortlist.
- A snapshot rotates and context fields update.
- A reviewer needs the side-by-side reference.
Review before moving on
- Comparison URL is captured.
- I named the spread without naming a winner.
- I read the citation behind at least one verified field.
- I planned a workload-specific test before integration.
Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.
What this exercise does not produce
- A model recommendation. The exercise routes you through evidence — you decide.
- A score or grade for any model. The catalogue does not score.
- A substitute for external prompt, latency, rate-limit, cost, or compliance tests.