Find
Resource finder
Find the lesson, exercise, lab playbook, prompt set, kit, or evidence workflow that matches your role and task. Every link below opens an existing surface — the finder routes you in, it does not recommend a model.
Total resources
62
Lessons, exercises, paths, lab tools, kits, outcomes, audiences, demos, workspaces, evidence examples.
Learn stage
19
Open filtered view →
Apply stage
7
Open filtered view →
Verify stage
7
Open filtered view →
Test stage
15
Open filtered view →
Package stage
14
Open filtered view →
Next step
I want to…
Each card opens a filtered view of the resource finder — the canonical URL stays /resources and filtered URLs are noindex,follow.
I want to learn the basics
Plain-language concept lessons.
Open filtered view →
I want to choose model candidates
Build a source-backed shortlist.
Open filtered view →
I want to compare models side by side
Render verified fields against each other.
Open filtered view →
I want to test model behaviour
Run prompt + structured-output + regression tests.
Open filtered view →
I want to evaluate prompts
Six generic, safe evaluation prompt sets.
Open filtered view →
I want to document evidence
Package the decision brief or evaluation plan.
Open filtered view →
I want to review sources
Audit citations + freshness across the catalogue.
Open filtered view →
I want to prepare a governance review
Source freshness + lifecycle + refusal-boundary suite.
Open filtered view →
I want to test an automation workflow
Validate model behaviour inside an unattended loop.
Open filtered view →
Stage map
Learn → Apply → Verify → Test → Package
Every product surface lives at exactly one stage. The graph counts the resources at each stage so the reader can scan where the next step lives.
Step 1
Learn
Plain-language concept lessons + audience entry points + role-based learning paths.
Step 2
Apply
Exercises + selection / comparison workspaces + guided demos that produce a working URL.
Step 3
Verify
Source freshness + lifecycle inspection + reverification queue + coverage audit.
Step 4
Test
Lab playbooks + evaluation prompt sets the reader runs in their own harness.
Step 5
Package
Decision brief + Markdown templates + workflow kits + outcome flows that ship a paste-ready artifact.
Filters
Every filter is a link — no client state, no accounts, no progress tracking. Filtered pages are noindex,follow; the canonical URL is /resources.
Goal
Resource type
Evidence artifact
Difficulty
Results
62 resources
Showing every resource in the graph. Apply filters above to narrow the view.
Learn · 19 resources
Lesson · Learn
How to choose an AI model
A workflow for picking which AI model to test next — start from your use case, inspect verified fields, export an evidence brief.
BeginnersDevelopersProduct teamsOpen →
Lesson · Learn
Context windows explained
What a context window means, what it does not guarantee, and which verified fields to inspect before assuming a model fits your prompt.
BeginnersDevelopersOpen →
Lesson · Learn
Hosted vs first-party AI models
Why the model creator and the billing provider are usually different, and how the catalogue keeps the two separate.
BeginnersDevelopersProduct teamsHosted-provider noteOpen →
Lesson · Learn
AI model pricing references explained
Why catalogue pricing rows are references, not quotes — and how to read them without ranking models by price.
Product teamsAutomation specialistsOpen →
Lesson · Learn
Model lifecycle: active, deprecated, retired
What active, preview, deprecated, and retired mean for a model — and why lifecycle should gate integration decisions.
Product teamsGovernance teamsLifecycle review noteOpen →
Lesson · Learn
How to test an AI model before integration
After the shortlist: how to run your own prompt, latency, rate-limit, cost, and compliance tests — using the evidence brief as the pack you ship to reviewers.
DevelopersAutomation specialistsExternal test planOpen →
Lesson · Learn
Multimodal input: image, audio, video, PDF
How the catalogue records which models accept image, audio, video, or PDF input — and why marketing copy is not enough to assume support.
DevelopersProduct teamsOpen →
Lesson · Learn
Structured output, JSON mode, and tool use
The difference between structured output, JSON mode, and tool/function calling — and what is currently verified in the catalogue.
DevelopersAutomation specialistsOpen →
Lesson · Learn
Status-aware model selection
Why vendor-reported status pages and independent probes are kept separate — and when status should gate a model decision.
Governance teamsAutomation specialistsOpen →
Lesson · Learn
Why benchmark scores can mislead
Contamination, prompt variance, version drift, and why the catalogue does not publish provider-reported benchmark scores casually.
BeginnersDevelopersProduct teamsOpen →
Learning path · Learn
AI model basics for careful users
Three foundational readings + four practical exercises. Walks from a use case to a paste-ready evidence brief plus a freshness checklist. Built for readers new to the catalogue.
BeginnersOpen →
Learning path · Learn
Technical model evaluation before integration
Four readings + three exercises + two pre-seeded workflows. Walks the verified fields a developer needs (hosted creator vs host, modality channels, structured generation, the testing framework) and ends with a comparison URL, an evidence brief, and a written external test plan.
DevelopersDecision briefOpen →
Learning path · Learn
Model selection for product use cases
Four readings + three exercises + one pre-seeded workflow. Walks use-case framing, pricing references, lifecycle gates, and benchmark limits so the team can align on a defensible review — ending with a use-case shortlist, a pricing-reference note, a lifecycle risk note, and the evidence brief.
Product teamsDecision briefOpen →
Learning path · Learn
AI model governance and source review
Four readings + three exercises + three audit workflows. Walks lifecycle, status, benchmark limits, pricing, and source freshness so a reviewer can sign off with a defensible evidence trail — never a certification claim.
Governance teamsSource freshness checklistLifecycle review noteOpen →
Learning path · Learn
Safe AI model use for automation workflows
Five readings + four exercises + three pre-seeded workflows. Built for people wiring AI models into automations: structured outputs, prompt cost projections, regression test plans, and a brief that ships with the automation runbook. Never an automation marketing pitch.
Automation specialistsAutomation risk checklistExternal test planOpen →
Audience · Learn
For developers
Evaluate AI models the way you evaluate any other infrastructure
DevelopersOpen →
Audience · Learn
For product teams
Turn a product use case into a defensible model decision
Product teamsOpen →
Audience · Learn
For automation specialists
Use AI models inside automations without over-trusting them
Automation specialistsOpen →
Audience · Learn
For governance teams
Build a defensible AI model review with sourced evidence
Governance teamsOpen →
Apply · 7 resources
Exercise · Apply
Build your first source-backed shortlist
Pick a use case, filter the catalogue by verified fields, and end with a shortlist URL you can share with the team.
BeginnersDevelopersProduct teamsShortlist URLOpen →
Exercise · Apply
Compare context windows without ranking models
Use the comparison builder to render verified context window + max output tokens for 3–4 candidate models side by side.
DevelopersProduct teamsComparison URLOpen →
Exercise · Apply
Map a hosted provider relationship
Pick a hosted model in the catalogue, trace creator vs billing provider, and read the hosted pricing reference's source citation.
DevelopersProduct teamsHosted-provider noteOpen →
Guided demo · Apply
Long-context analysis
Walk a long-context (≥200k-token) workload: pick the use case, narrow a shortlist by verified context window, compare the candidates side by side, and export an evidence brief. Pricing tier references for prompts >200k are flagged on the comparison; the brief surfaces every data gap explicitly.
DevelopersProduct teamsComparison URLShortlist URLOpen →
Guided demo · Apply
Hosted inference
Walk a hosted-inference workflow where the model creator does not run a paid first-party API. Pick the use case, narrow the shortlist to models with verified hosted availability, then compare hosted pricing references side by side. The brief separates the hosting platform (billing provider) from the model creator at every step.
DevelopersProduct teamsHosted-provider noteShortlist URLOpen →
Workspace · Apply
Selection workspace
Narrow a source-backed shortlist using verified catalogue fields. Output is a /select URL anyone can re-open.
BeginnersDevelopersProduct teamsShortlist URLOpen →
Workspace · Apply
Comparison builder
Render up to four candidate models side by side against verified fields. Output is a /compare/build URL.
DevelopersProduct teamsComparison URLOpen →
Verify · 7 resources
Exercise · Verify
Review a pricing reference safely
Open a verified pricing row, read its unit semantics + retrieval date, and walk the reverification queue if it is stale.
Product teamsAutomation specialistsOpen →
Exercise · Verify
Inspect lifecycle before integration
Pull lifecycle state for a candidate model, check for retirement date, and add a migration target to your notes if one exists.
Product teamsGovernance teamsLifecycle review noteOpen →
Exercise · Verify
Check source freshness and reverification state
Open the sources hub for a provider, identify any stale citations, and walk the reverification queue to see what is due for re-check.
Governance teamsSource freshness checklistOpen →
Guided demo · Verify
Governance review
Walk an internal AI-inventory + source-backed due-diligence flow. Pick the use case, narrow a shortlist with verified citations, compare verification state side by side, then generate an evidence brief that lists every data gap and stale citation explicitly. The reverification queue is the natural next stop.
Governance teamsSource freshness checklistOpen →
Reference · Verify
Citation registry
Every primary-source citation backing a verified value in the catalogue. Filter by provider or source type.
Governance teamsProduct teamsOpen →
Reference · Verify
Coverage audit
Per-provider verified-field counts + citation density across the entity graph.
Governance teamsProduct teamsOpen →
Reference · Verify
Reverification queue
Citations due for manual re-check, with the date each source was last verified.
Governance teamsSource freshness checklistOpen →
Test · 15 resources
Exercise · Test
Plan an external model test
Use the brief and the testing lesson to plan your own prompt, latency, rate-limit, and cost validation work for the shortlist.
DevelopersAutomation specialistsGovernance teamsExternal test planOpen →
Lab playbook · Test
Prompt testing basics
The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.
DevelopersProduct teamsAutomation specialistsPrompt test matrixOpen →
Lab playbook · Test
Structured output testing
How to validate JSON mode, structured output, and tool calls against your real schema before depending on the model in a pipeline.
DevelopersAutomation specialistsPrompt test matrixOpen →
Lab playbook · Test
Long-context testing
How to test long-prompt behaviour past the catalogue's verified context window — recall, instruction adherence, and cost growth — without trusting a marketing number.
DevelopersProduct teamsPrompt test matrixOpen →
Lab playbook · Test
Multimodal input testing
How to test image, audio, video, and PDF input channels against your real assets — never against marketing copy.
DevelopersProduct teamsPrompt test matrixOpen →
Lab playbook · Test
Automation workflow testing
How to test a model inside an automation loop — chained prompts, retries, downstream parsers, regression surface — before letting it run unattended.
Automation specialistsAutomation risk checklistExternal test planOpen →
Lab playbook · Test
Model regression testing
How to run a small, repeatable canary suite after every snapshot rotation so silent regressions surface before production traffic notices.
Automation specialistsGovernance teamsPrompt test matrixExternal test planOpen →
Evaluation prompt set · Test
Summarization quality
Evaluate whether a model summarises without adding unsupported claims, omitting constraints, or inventing numbers.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Structured extraction
Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.
DevelopersAutomation specialistsPrompt test matrixOpen →
Evaluation prompt set · Test
Long-context recall
Evaluate whether a model preserves constraints, handles cross-references, and detects conflicts across multiple sections of input.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Instruction following
Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.
DevelopersProduct teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Refusal boundary
Evaluate whether the model handles benign boundary-setting safely — without over-refusing, without giving definitive professional advice, and without complying with inappropriate requests.
Governance teamsPrompt test matrixOpen →
Evaluation prompt set · Test
Automation robustness
Evaluate whether a model handles automation-style constraints (allowed categories, missing values, retry decisions, ambiguity flags) without silently breaking the contract.
Automation specialistsPrompt test matrixOpen →
Outcome · Test
LLM prompt evaluation
Evaluate model behaviour with safe, structured prompt sets
DevelopersProduct teamsPrompt test matrixOpen →
Outcome · Test
Structured output testing
Validate JSON mode, structured output, and tool calling against your real schema
DevelopersAutomation specialistsPrompt test matrixOpen →
Package · 14 resources
Exercise · Package
Create a decision evidence brief
Use the decision brief builder to generate a paste-ready evidence pack from your shortlist, then export it in Markdown.
DevelopersProduct teamsGovernance teamsDecision briefOpen →
Lab template · Package
Model evaluation plan
A blank evaluation plan for a single model + workload pairing. Use one copy per candidate.
DevelopersProduct teamsModel evaluation planOpen →
Lab template · Package
Prompt test matrix
A row-per-prompt matrix you fill in per candidate model. Pair with the model evaluation plan.
DevelopersProduct teamsAutomation specialistsPrompt test matrixOpen →
Lab template · Package
Automation risk checklist
A pre-launch risk checklist for automations that depend on a model. Pair with the automation workflow testing playbook.
Automation specialistsGovernance teamsAutomation risk checklistOpen →
Workflow kit · Package
Developer model evaluation kit
Prepare a source-backed model evaluation plan before integration. Walks the developer learning path, the matching exercises, the prompt-testing + structured-output playbooks, the structured-extraction + instruction-following prompt sets, and the model evaluation plan + prompt test matrix templates.
DevelopersDecision briefModel evaluation planPrompt test matrixOpen →
Workflow kit · Package
Automation workflow testing kit
Prepare a safe testing workflow for AI-powered automation. Walks the automation-specialist learning path, the structured-output + pricing-references + testing lessons, three exercises, the automation workflow testing + regression playbooks, the automation-robustness + structured-extraction prompt sets, and the automation risk checklist + prompt test matrix templates.
Automation specialistsAutomation risk checklistPrompt test matrixExternal test planOpen →
Workflow kit · Package
Product model selection kit
Turn a product use case into a reviewable model selection artifact. Walks the product manager learning path, four lessons covering use-case framing through benchmark limits, three exercises that produce pricing + lifecycle notes + the brief, the prompt-testing + long-context playbooks, two prompt sets, and the model evaluation plan template.
Product teamsDecision briefShortlist URLOpen →
Workflow kit · Package
Governance review kit
Prepare a source / freshness / lifecycle review package for internal governance discussions. Walks the governance learning path, four lessons (lifecycle, status, benchmark limits, pricing), three exercises that produce the freshness checklist + lifecycle note + test plan, the regression + prompt-testing playbooks, the refusal-boundary + instruction-following prompt sets, and the evaluation plan + automation risk checklist templates.
Governance teamsDecision briefSource freshness checklistLifecycle review noteOpen →
Outcome · Package
AI model evaluation for developers
Evaluate AI models before integration without trusting a leaderboard
DevelopersDecision briefExternal test planOpen →
Outcome · Package
AI model selection for product teams
Turn a product use case into a defensible model selection
Product teamsDecision briefOpen →
Outcome · Package
AI automation testing
Test an AI model inside an automation before it runs unattended
Automation specialistsAutomation risk checklistExternal test planOpen →
Outcome · Package
AI model governance review
Build a defensible governance review with sourced evidence
Governance teamsSource freshness checklistLifecycle review noteOpen →
Workspace · Package
Decision brief builder
Build a paste-ready Markdown evidence brief from the verified catalogue. Output is a /briefs/build URL plus the Markdown.
DevelopersProduct teamsGovernance teamsDecision briefOpen →
Evidence example · Package
Example decision brief
Worked example produced by the same buildDecisionBrief() helper as the live builder.
DevelopersProduct teamsGovernance teamsDecision briefOpen →