Learn · Path · Engineer preparing an integration
Technical model evaluation before integration
Four readings + three exercises + two pre-seeded workflows. Walks the verified fields a developer needs (hosted creator vs host, modality channels, structured generation, the testing framework) and ends with a comparison URL, an evidence brief, and a written external test plan.
Use this path inside a kit
The matching workflow kit bundles this path with the lab playbooks, evaluation prompt sets, and Markdown templates the role needs, and exports the whole thing as a single work document.
Audience
Engineer preparing an integration
Difficulty
intermediate
Estimated time
112 min · 11 steps
What you walk away with
What you will learn
- How the catalogue separates model creator from billing provider.
- Why marketing copy is not enough to assume modality support.
- The difference between JSON mode, structured output, and tool calling.
- Which workload-specific tests the catalogue cannot run for you.
What you will build
- A hosted-provider mapping note (creator + host + data gap).
- A pre-seeded /compare/build URL for hosted inference.
- A Markdown evidence brief covering the technical evaluation fields.
- A written external test plan paired with the brief.
Evidence artifacts
- Hosted-provider mapping note
- Verified comparison URL
- Decision evidence brief
- External test plan
Tools used
Learn hub →
Concept lessons that explain each verified field.
Exercises →
Practical workflows that end with evidence artifacts.
Comparison builder →
Render verified fields side by side for 2–4 models.
Decision brief builder →
Export a Markdown or JSON evidence brief.
Citation registry →
Every primary-source URL the catalogue references.
Coverage audit →
Per-provider verified-field counts and citation density.
Prerequisites
- You have read /learn/how-to-choose-ai-model OR walked the beginner path.
- You can name the model you are evaluating — even tentatively.
Timeline
11 steps
Open each step in the order shown. The route on every step is the canonical workspace or lesson — no parallel UI.
- lesson5 min
Hosted vs first-party AI models
Separate the model creator from the billing provider before you integrate.
- lesson5 min
Multimodal input: image, audio, video, PDF
Confirm the input channels the catalogue actually verifies.
- lesson5 min
Structured output, JSON mode, and tool use
Distinguish JSON mode from structured output from tool calling.
- lesson5 min
How to test an AI model before integration
Read the testing framework the catalogue does not run for you.
- exercise10 min
Map a hosted provider relationship
End with a creator + billing platform + data-gap note for a hosted model.
- exercise10 min
Create a decision evidence brief
End with a Markdown brief covering the technical fields your reviewer needs.
- exercise12 min
Plan an external model test
End with a written test plan that pairs the brief with workload-specific tests.
- workflow5 min
Open the hosted-inference comparison builder
Pre-seed the builder so the verified hosted-availability + hosted-pricing fields land in your evidence.
- workflow5 min
Open the hosted-inference brief builder
Pre-seed the brief builder so the evidence pack stays consistent with the comparison.
- workflow25 min
Run the prompt-testing playbook
Walk a minimal repeatable prompt-test routine against the shortlisted candidates and end with a Markdown evidence brief.
- workflow25 min
Run the structured-extraction prompt set
Test schema-conformant extraction against your real schema before integration.
Start next
Step 1 of 11: Hosted vs first-party AI models
Separate the model creator from the billing provider before you integrate.
What this path does not promise
- A recommended hosting platform.
- Asserted latency, throughput, or uptime numbers.
- Production readiness without your own tests.
How to use this path
- Open each route in the order shown. The path is a sequenced reading + practice plan, nothing more.
- Keep the artifacts you produce — the shortlist URL, comparison URL, Markdown brief, freshness checklist, or test plan.
- There is no login surface, no progress state, and no completion certificate.
- External workload-specific testing remains your team's responsibility — the catalogue surfaces evidence, not verdicts.
No progress, no accounts, no certificates
- No accounts — the catalogue does not have a login surface.
- No progress tracking — the catalogue does not store which pages you have visited.
- No certificates — the catalogue does not issue completion credentials, badges, or scores.
- Completion is the artifact you produce: a shortlist URL, a comparison URL, a Markdown brief, a freshness checklist, or a written test plan.