Skip to content
WebmasterID

Exercise · beginner

Build your first source-backed shortlist

Pick a use case, filter the catalogue by verified fields, and end with a shortlist URL you can share with the team.

Estimated time: 8 minutes. Exercises produce evidence artifacts, never model recommendations.

Goal

Produce a deterministic shortlist URL that captures which use case, which verified-field filters, and which model order the catalogue surfaced.

Prerequisites

  • Read /learn/how-to-choose-ai-model so the workflow framing makes sense.
  • Know which workload you are evaluating (long-context, multimodal, hosted, governance, or another existing use case).

Step-by-step exercise

  1. Pick a use case

    Open the use cases hub and choose the one that matches the workload you are evaluating.

    Open /use-cases →

    Expected outcome: You can name which verified fields the use case asks the catalogue to weight.

  2. Open the selection workspace with your use case pre-selected

    Visit /select and set the Use case filter, then narrow further by lifecycle (active only) and verified pricing coverage if relevant.

    Open /select →

    Expected outcome: The URL bar now carries your filter selections — copy this URL, it is your shortlist artifact.

  3. Read the deterministic order

    The shortlist orders by verified field count, active lifecycle, source count, name — never by score. Read the rows top to bottom.

    Open /select →

    Expected outcome: You can describe out loud why each row is ordered where it is, without invoking quality or price as the reason.

  4. Open a top-three candidate model page

    Click any model name in the shortlist. Read the verified-field rows and the citation count.

    Open /models →

    Expected outcome: You can name at least one data gap (an unverified field) for that model.

Completion checklist

Completion checklist

  • The shortlist URL is saved.
  • You can describe why the shortlist ordered itself the way it did.
  • You named at least one data gap on the top candidate.
  • You did NOT pick a winner.

Evidence artifact

At the end of this exercise you should have: A /select URL with your use-case + filter parameters that opens the same shortlist for any teammate.

Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.

Related workflow routes

Exercise does not recommend a model — external testing still required.

Common mistake

Treating the first row of the shortlist as the answer. Shortlist order is deterministic (verified field count → lifecycle → source count → name), not a quality ranking.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Example artifact

## Shortlist URL (illustrative)
/select?useCase=<your-use-case>&lifecycle=active&pricingCoverage=verified

Top 3 rows (in catalogue order):
- <candidate-A> · verified context: <tokens> · active
- <candidate-B> · verified context: <tokens> · active
- <candidate-C> · verified context: <tokens> · active

Data gap noted: <candidate-C> has unverified max-output. Plan to confirm externally.

Substitute your real shortlist, slugs, dates, and values when you walk the exercise.

Repeat this exercise when

  • The use case or workload changes substantially.
  • A snapshot rotates and lifecycle fields shift.
  • Pricing references go stale and need re-verification.
  • A teammate needs to inherit the shortlist URL.

Review before moving on

  • Shortlist URL is captured (copy from the address bar).
  • I can describe ordering rationale without using quality language.
  • At least one data gap is named on the top candidate.
  • No winner is declared in the notes.

Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.

What this exercise does not produce

  • A model recommendation. The exercise routes you through evidence — you decide.
  • A score or grade for any model. The catalogue does not score.
  • A substitute for external prompt, latency, rate-limit, cost, or compliance tests.