Skip to content
WebmasterID

Exercise · beginner

Compare context windows without ranking models

Use the comparison builder to render verified context window + max output tokens for 3–4 candidate models side by side.

Estimated time: 7 minutes. Exercises produce evidence artifacts, never model recommendations.

Goal

Render a side-by-side comparison page where the only verified fields you look at are context window and max output tokens — and articulate what those values do not tell you.

Prerequisites

  • Read /learn/context-window so the difference between context window and max output is clear.
  • Have a list of 3–4 candidate model slugs in mind (your shortlist from the previous exercise, or any active models from /models).

Step-by-step exercise

  1. Open the comparison builder

    Visit /compare/build and select 3 or 4 candidate models. Optionally set the use-case filter.

    Open /compare/build →

    Expected outcome: The page renders columns for your selected models. The URL captures the selection.

  2. Read context window and max output rows

    Locate the verified context window and max output token rows. Note which columns render a verified value vs the unverified-data label.

    Open /compare/build →

    Expected outcome: You can describe the spread of context windows and max output limits across the columns — without naming a winner.

  3. Open the citation for the largest context window

    Click through to the model page for the row with the largest verified context window. Read the citation it points at.

    Open /models →

    Expected outcome: You can name the primary source the catalogue cited for that value and when it was retrieved.

  4. Note what context size does not guarantee

    Write down (or paste into your notes) the workload-specific behaviours the catalogue does not measure: deep-recall accuracy, instruction-following at length, prompt-size pricing tier impact.

    Open /learn/context-window →

    Expected outcome: Your notes capture at least three things you still need to test in your own environment.

Completion checklist

Completion checklist

  • The comparison URL is saved.
  • You can describe the spread of context windows.
  • You read the source URL for at least one verified value.
  • You did NOT pick the model with the largest context window as the winner.

Evidence artifact

At the end of this exercise you should have: A /compare/build URL with 3–4 model slugs that any teammate can open to see the same comparison.

Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.

Related workflow routes

Exercise does not recommend a model — external testing still required.

Common mistake

Treating the candidate with the largest verified context window as the answer. Larger context does not equal better recall or better fit.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Example artifact

## Context comparison URL (illustrative)
/compare/build?models=<slug-A>,<slug-B>,<slug-C>&useCase=long-context-analysis

Spread observed:
- Context window: <range>
- Max output: <range>

Notes to capture before testing:
- Recall behaviour is workload-specific.
- Pricing tiers may scale non-linearly with prompt size.

Substitute your real shortlist, slugs, dates, and values when you walk the exercise.

Repeat this exercise when

  • A new candidate is added to the shortlist.
  • A snapshot rotates and context fields update.
  • A reviewer needs the side-by-side reference.

Review before moving on

  • Comparison URL is captured.
  • I named the spread without naming a winner.
  • I read the citation behind at least one verified field.
  • I planned a workload-specific test before integration.

Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.

What this exercise does not produce

  • A model recommendation. The exercise routes you through evidence — you decide.
  • A score or grade for any model. The catalogue does not score.
  • A substitute for external prompt, latency, rate-limit, cost, or compliance tests.