Kit · developers
Developer model evaluation kit
Prepare a source-backed model evaluation plan before integration. Walks the developer learning path, the matching exercises, the prompt-testing + structured-output playbooks, the structured-extraction + instruction-following prompt sets, and the model evaluation plan + prompt test matrix templates.
Audience
developers
Difficulty
intermediate
Estimated time
180 min · 8 steps
Who this kit is for
Verified model fields, structured testing playbooks, and Markdown-exportable evidence briefs for engineers preparing an integration. The platform never declares a winner — you decide which candidate fits your workload.
Goal
End with a Markdown evidence brief plus a written external test plan that pairs verified catalogue fields with workload-specific tests.
What you will produce
- Hosted-provider mapping note
- Comparison URL from /compare/build
- Model evaluation plan (paste-ready Markdown)
- Prompt test matrix with per-candidate observations
- Decision evidence brief
- External test plan
Prerequisites
Export Markdown
The kit serialises to a single Markdown document you can paste into a design doc, ticket, or PR description.
Workflow
8 sequenced steps
Open each step in order. Every step opens a route that already exists — no parallel UI.
Step-by-step workflow
Walk the developer learning path
Read the four lessons in order so the verified fields the rest of the kit depends on are framed.
Output: Notes on hosted vs first-party, modality, structured output, testing framework.
Map the hosted provider
Complete the map-hosted-provider exercise. Separate model creator from billing platform per candidate.
Open
/learn/exercises/map-hosted-provider→Output: Hosted-provider mapping note (creator + host + data gap).
Build the side-by-side comparison
Open /compare/build with the candidate slugs and copy the URL into your notes.
Output: Comparison URL that opens the same view for any teammate.
Run the prompt-testing playbook
Walk the minimum prompt-testing routine against the candidates in your own harness.
Open
/lab/prompt-testing-basics→Output: Per-prompt observations recorded against the acceptance rubric.
Run the structured-extraction prompt set
Open the structured-extraction prompt set and run it against your real schema in your harness.
Open
/lab/prompts/structured-extraction→Output: Per-prompt schema validity record + raw responses.
Fill in the prompt test matrix template
Paste your per-candidate observations into the prompt-test-matrix template.
Open
/lab/templates/prompt-test-matrix→Output: Markdown matrix attached to the brief.
Generate the decision evidence brief
Open /briefs/build with the candidate slugs and export Markdown.
Output: Markdown brief listing verified fields, data gaps, source trail, freshness.
Write the external test plan
Complete the plan-external-model-test exercise. Pair the brief with workload-specific tests.
Open
/learn/exercises/plan-external-model-test→Output: Written test plan covering prompts, latency, rate limits, cost, compliance, regression.
Required resources
The kit reuses existing surfaces — no parallel UI, no duplicated content. Open each surface in the order the timeline lists.
Lessons
Hosted vs first-party AI models →
Why the model creator and the billing provider are usually different, and how the catalogue keeps the two separate.
Structured output, JSON mode, and tool use →
The difference between structured output, JSON mode, and tool/function calling — and what is currently verified in the catalogue.
How to test an AI model before integration →
After the shortlist: how to run your own prompt, latency, rate-limit, cost, and compliance tests — using the evidence brief as the pack you ship to reviewers.
Exercises
Map a hosted provider relationship →
Pick a hosted model in the catalogue, trace creator vs billing provider, and read the hosted pricing reference's source citation.
Create a decision evidence brief →
Use the decision brief builder to generate a paste-ready evidence pack from your shortlist, then export it in Markdown.
Plan an external model test →
Use the brief and the testing lesson to plan your own prompt, latency, rate-limit, and cost validation work for the shortlist.
Lab playbooks
Prompt testing basics →
The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.
Structured output testing →
How to validate JSON mode, structured output, and tool calls against your real schema before depending on the model in a pipeline.
Evaluation prompt sets
Structured extraction →
Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.
Instruction following →
Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.
Final checklist
- Hosted-provider mapping note captured (creator + host + data gap).
- Comparison URL saved.
- Prompt-testing observations recorded per prompt, not as a single score.
- Schema validity record captured for the structured-extraction set.
- Decision evidence brief exported in Markdown.
- External test plan written, with regression cadence named.
Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.
Evidence routes
What this kit does not promise
- Pick the right model for your integration.
- Assert latency, throughput, or uptime.
- Certify the model for any regulatory regime.
- Replace your own workload-specific testing.
What workflow kits do not promise
- No model recommendations or rankings. Kits walk the evidence; the reader's team makes the decision.
- No live pricing quotes. Pricing rows referenced inside the kit are sourced references with retrievedAt dates.
- No production-readiness guarantee. The kit ends with an external test plan — running those tests is the team's responsibility.
- No compliance certification, legal advice, or vendor endorsement.
- No SEO ranking guarantees or automation reliability guarantees.
- No accounts, no progress tracking, no course-completion certificates. Completion is the Markdown artifacts the kit puts in your hands.