Find
Resource finder
Find the lesson, exercise, lab playbook, prompt set, kit, or evidence workflow that matches your role and task. Every link below opens an existing surface — the finder routes you in, it does not recommend a model.
Total resources
62
Lessons, exercises, paths, lab tools, kits, outcomes, audiences, demos, workspaces, evidence examples.
Learn stage
19
Open filtered view →
Apply stage
7
Open filtered view →
Verify stage
7
Open filtered view →
Test stage
15
Open filtered view →
Package stage
14
Open filtered view →
Next step
I want to…
Each card opens a filtered view of the resource finder — the canonical URL stays /resources and filtered URLs are noindex,follow.
I want to learn the basics
Plain-language concept lessons.
Open filtered view →
I want to choose model candidates
Build a source-backed shortlist.
Open filtered view →
I want to compare models side by side
Render verified fields against each other.
Open filtered view →
I want to test model behaviour
Run prompt + structured-output + regression tests.
Open filtered view →
I want to evaluate prompts
Six generic, safe evaluation prompt sets.
Open filtered view →
I want to document evidence
Package the decision brief or evaluation plan.
Open filtered view →
I want to review sources
Audit citations + freshness across the catalogue.
Open filtered view →
I want to prepare a governance review
Source freshness + lifecycle + refusal-boundary suite.
Open filtered view →
I want to test an automation workflow
Validate model behaviour inside an unattended loop.
Open filtered view →
Stage map
Learn → Apply → Verify → Test → Package
Every product surface lives at exactly one stage. The graph counts the resources at each stage so the reader can scan where the next step lives.
Step 1
Learn
Plain-language concept lessons + audience entry points + role-based learning paths.
Step 2
Apply
Exercises + selection / comparison workspaces + guided demos that produce a working URL.
Step 3
Verify
Source freshness + lifecycle inspection + reverification queue + coverage audit.
Step 4
Test
Lab playbooks + evaluation prompt sets the reader runs in their own harness.
Step 5
Package
Decision brief + Markdown templates + workflow kits + outcome flows that ship a paste-ready artifact.
Filters
Reset all filters →Every filter is a link — no client state, no accounts, no progress tracking. Filtered pages are noindex,follow; the canonical URL is /resources.
Goal
Resource type
Evidence artifact
Difficulty
Results
34 resources
Filtered view: audience: Developers. Canonical URL stays /resources; this filtered URL is noindex,follow.
Learn · 9 resources
Lesson · Learn
How to choose an AI model
A workflow for picking which AI model to test next — start from your use case, inspect verified fields, export an evidence brief.
BeginnersDevelopersProduct teamsOpen →
Lesson · Learn
Context windows explained
What a context window means, what it does not guarantee, and which verified fields to inspect before assuming a model fits your prompt.
BeginnersDevelopersOpen →
Lesson · Learn
Hosted vs first-party AI models
Why the model creator and the billing provider are usually different, and how the catalogue keeps the two separate.
BeginnersDevelopersProduct teamsHosted-provider noteOpen →
Lesson · Learn
How to test an AI model before integration
After the shortlist: how to run your own prompt, latency, rate-limit, cost, and compliance tests — using the evidence brief as the pack you ship to reviewers.
DevelopersAutomation specialistsExternal test planOpen →
Lesson · Learn
Multimodal input: image, audio, video, PDF
How the catalogue records which models accept image, audio, video, or PDF input — and why marketing copy is not enough to assume support.
DevelopersProduct teamsOpen →
Lesson · Learn
Structured output, JSON mode, and tool use
The difference between structured output, JSON mode, and tool/function calling — and what is currently verified in the catalogue.
DevelopersAutomation specialistsOpen →
Lesson · Learn
Why benchmark scores can mislead
Contamination, prompt variance, version drift, and why the catalogue does not publish provider-reported benchmark scores casually.
BeginnersDevelopersProduct teamsOpen →
Learning path · Learn
Technical model evaluation before integration
Four readings + three exercises + two pre-seeded workflows. Walks the verified fields a developer needs (hosted creator vs host, modality channels, structured generation, the testing framework) and ends with a comparison URL, an evidence brief, and a written external test plan.
DevelopersDecision briefOpen →
Audience · Learn
For developers
Evaluate AI models the way you evaluate any other infrastructure
DevelopersOpen →
Apply · 7 resources
Exercise · Apply
Build your first source-backed shortlist
Pick a use case, filter the catalogue by verified fields, and end with a shortlist URL you can share with the team.
BeginnersDevelopersProduct teamsShortlist URLOpen →
Exercise · Apply
Compare context windows without ranking models
Use the comparison builder to render verified context window + max output tokens for 3–4 candidate models side by side.
DevelopersProduct teamsComparison URLOpen →
Exercise · Apply
Map a hosted provider relationship
Pick a hosted model in the catalogue, trace creator vs billing provider, and read the hosted pricing reference's source citation.
DevelopersProduct teamsHosted-provider noteOpen →
Guided demo · Apply
Long-context analysis
Walk a long-context (≥200k-token) workload: pick the use case, narrow a shortlist by verified context window, compare the candidates side by side, and export an evidence brief. Pricing tier references for prompts >200k are flagged on the comparison; the brief surfaces every data gap explicitly.
DevelopersProduct teamsComparison URLShortlist URLOpen →
Guided demo · Apply
Hosted inference
Walk a hosted-inference workflow where the model creator does not run a paid first-party API. Pick the use case, narrow the shortlist to models with verified hosted availability, then compare hosted pricing references side by side. The brief separates the hosting platform (billing provider) from the model creator at every step.
DevelopersProduct teamsHosted-provider noteShortlist URLOpen →
Workspace · Apply
Selection workspace
Narrow a source-backed shortlist using verified catalogue fields. Output is a /select URL anyone can re-open.
BeginnersDevelopersProduct teamsShortlist URLOpen →
Workspace · Apply
Comparison builder
Render up to four candidate models side by side against verified fields. Output is a /compare/build URL.
DevelopersProduct teamsComparison URLOpen →
Test · 11 resources
Exercise · Test
Plan an external model test
Use the brief and the testing lesson to plan your own prompt, latency, rate-limit, and cost validation work for the shortlist.
DevelopersAutomation specialistsGovernance teamsExternal test planOpen →
Lab playbook · Test
Prompt testing basics
The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.
DevelopersProduct teamsAutomation specialistsPrompt test matrixOpen →
Lab playbook · Test
Structured output testing
How to validate JSON mode, structured output, and tool calls against your real schema before depending on the model in a pipeline.
DevelopersAutomation specialistsPrompt test matrixOpen →
Lab playbook · Test
Long-context testing
How to test long-prompt behaviour past the catalogue's verified context window — recall, instruction adherence, and cost growth — without trusting a marketing number.
DevelopersProduct teamsPrompt test matrixOpen →
Lab playbook · Test
Multimodal input testing
How to test image, audio, video, and PDF input channels against your real assets — never against marketing copy.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Summarization quality
Evaluate whether a model summarises without adding unsupported claims, omitting constraints, or inventing numbers.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Structured extraction
Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.
DevelopersAutomation specialistsPrompt test matrixOpen →
Evaluation prompt set · Test
Long-context recall
Evaluate whether a model preserves constraints, handles cross-references, and detects conflicts across multiple sections of input.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Instruction following
Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.
DevelopersProduct teamsPrompt test matrixOpen →
Outcome · Test
LLM prompt evaluation
Evaluate model behaviour with safe, structured prompt sets
DevelopersProduct teamsPrompt test matrixOpen →
Outcome · Test
Structured output testing
Validate JSON mode, structured output, and tool calling against your real schema
DevelopersAutomation specialistsPrompt test matrixOpen →
Package · 7 resources
Exercise · Package
Create a decision evidence brief
Use the decision brief builder to generate a paste-ready evidence pack from your shortlist, then export it in Markdown.
DevelopersProduct teamsGovernance teamsDecision briefOpen →
Lab template · Package
Model evaluation plan
A blank evaluation plan for a single model + workload pairing. Use one copy per candidate.
DevelopersProduct teamsModel evaluation planOpen →
Lab template · Package
Prompt test matrix
A row-per-prompt matrix you fill in per candidate model. Pair with the model evaluation plan.
DevelopersProduct teamsAutomation specialistsPrompt test matrixOpen →
Workflow kit · Package
Developer model evaluation kit
Prepare a source-backed model evaluation plan before integration. Walks the developer learning path, the matching exercises, the prompt-testing + structured-output playbooks, the structured-extraction + instruction-following prompt sets, and the model evaluation plan + prompt test matrix templates.
DevelopersDecision briefModel evaluation planPrompt test matrixOpen →
Outcome · Package
AI model evaluation for developers
Evaluate AI models before integration without trusting a leaderboard
DevelopersDecision briefExternal test planOpen →
Workspace · Package
Decision brief builder
Build a paste-ready Markdown evidence brief from the verified catalogue. Output is a /briefs/build URL plus the Markdown.
DevelopersProduct teamsGovernance teamsDecision briefOpen →
Evidence example · Package
Example decision brief
Worked example produced by the same buildDecisionBrief() helper as the live builder.
DevelopersProduct teamsGovernance teamsDecision briefOpen →