Skip to content
WebmasterID

Learn · Path · Engineer preparing an integration

Technical model evaluation before integration

Four readings + three exercises + two pre-seeded workflows. Walks the verified fields a developer needs (hosted creator vs host, modality channels, structured generation, the testing framework) and ends with a comparison URL, an evidence brief, and a written external test plan.

Use this path inside a kit

The matching workflow kit bundles this path with the lab playbooks, evaluation prompt sets, and Markdown templates the role needs, and exports the whole thing as a single work document.

Open the Developer model evaluation kit →

Audience

Engineer preparing an integration

Difficulty

intermediate

Estimated time

112 min · 11 steps

What you walk away with

What you will learn

  • How the catalogue separates model creator from billing provider.
  • Why marketing copy is not enough to assume modality support.
  • The difference between JSON mode, structured output, and tool calling.
  • Which workload-specific tests the catalogue cannot run for you.

What you will build

  • A hosted-provider mapping note (creator + host + data gap).
  • A pre-seeded /compare/build URL for hosted inference.
  • A Markdown evidence brief covering the technical evaluation fields.
  • A written external test plan paired with the brief.

Evidence artifacts

  • Hosted-provider mapping note
  • Verified comparison URL
  • Decision evidence brief
  • External test plan

Prerequisites

  • You have read /learn/how-to-choose-ai-model OR walked the beginner path.
  • You can name the model you are evaluating — even tentatively.

Timeline

11 steps

Open each step in the order shown. The route on every step is the canonical workspace or lesson — no parallel UI.

  1. lesson5 min

    Hosted vs first-party AI models

    Separate the model creator from the billing provider before you integrate.

    Open /learn/hosted-vs-first-party →

  2. lesson5 min

    Multimodal input: image, audio, video, PDF

    Confirm the input channels the catalogue actually verifies.

    Open /learn/multimodal-input →

  3. lesson5 min

    Structured output, JSON mode, and tool use

    Distinguish JSON mode from structured output from tool calling.

    Open /learn/structured-output →

  4. lesson5 min

    How to test an AI model before integration

    Read the testing framework the catalogue does not run for you.

    Open /learn/testing-ai-models →

  5. exercise10 min

    Map a hosted provider relationship

    End with a creator + billing platform + data-gap note for a hosted model.

    Open /learn/exercises/map-hosted-provider →

  6. exercise10 min

    Create a decision evidence brief

    End with a Markdown brief covering the technical fields your reviewer needs.

    Open /learn/exercises/create-decision-brief →

  7. exercise12 min

    Plan an external model test

    End with a written test plan that pairs the brief with workload-specific tests.

    Open /learn/exercises/plan-external-model-test →

  8. workflow5 min

    Open the hosted-inference comparison builder

    Pre-seed the builder so the verified hosted-availability + hosted-pricing fields land in your evidence.

    Open /compare/build?useCase=hosted-inference →

  9. workflow5 min

    Open the hosted-inference brief builder

    Pre-seed the brief builder so the evidence pack stays consistent with the comparison.

    Open /briefs/build?useCase=hosted-inference →

  10. workflow25 min

    Run the prompt-testing playbook

    Walk a minimal repeatable prompt-test routine against the shortlisted candidates and end with a Markdown evidence brief.

    Open /lab/prompt-testing-basics →

  11. workflow25 min

    Run the structured-extraction prompt set

    Test schema-conformant extraction against your real schema before integration.

    Open /lab/prompts/structured-extraction →

Start next

Step 1 of 11: Hosted vs first-party AI models

Separate the model creator from the billing provider before you integrate.

Open first step →

What this path does not promise

  • A recommended hosting platform.
  • Asserted latency, throughput, or uptime numbers.
  • Production readiness without your own tests.

How to use this path

  • Open each route in the order shown. The path is a sequenced reading + practice plan, nothing more.
  • Keep the artifacts you produce — the shortlist URL, comparison URL, Markdown brief, freshness checklist, or test plan.
  • There is no login surface, no progress state, and no completion certificate.
  • External workload-specific testing remains your team's responsibility — the catalogue surfaces evidence, not verdicts.

No progress, no accounts, no certificates

  • No accounts — the catalogue does not have a login surface.
  • No progress tracking — the catalogue does not store which pages you have visited.
  • No certificates — the catalogue does not issue completion credentials, badges, or scores.
  • Completion is the artifact you produce: a shortlist URL, a comparison URL, a Markdown brief, a freshness checklist, or a written test plan.