Skip to content
WebmasterID

Kit · developers

Developer model evaluation kit

Prepare a source-backed model evaluation plan before integration. Walks the developer learning path, the matching exercises, the prompt-testing + structured-output playbooks, the structured-extraction + instruction-following prompt sets, and the model evaluation plan + prompt test matrix templates.

Audience

developers

Difficulty

intermediate

Estimated time

180 min · 8 steps

Who this kit is for

Verified model fields, structured testing playbooks, and Markdown-exportable evidence briefs for engineers preparing an integration. The platform never declares a winner — you decide which candidate fits your workload.

Open the audience page → /for/developers

Goal

End with a Markdown evidence brief plus a written external test plan that pairs verified catalogue fields with workload-specific tests.

What you will produce

  • Hosted-provider mapping note
  • Comparison URL from /compare/build
  • Model evaluation plan (paste-ready Markdown)
  • Prompt test matrix with per-candidate observations
  • Decision evidence brief
  • External test plan

Prerequisites

Export Markdown

The kit serialises to a single Markdown document you can paste into a design doc, ticket, or PR description.

Open raw Markdown → /api/kits/developer-model-evaluation

Workflow

8 sequenced steps

Open each step in order. Every step opens a route that already exists — no parallel UI.

Step-by-step workflow

  1. Walk the developer learning path

    Read the four lessons in order so the verified fields the rest of the kit depends on are framed.

    Open /learn/path/developer →

    Output: Notes on hosted vs first-party, modality, structured output, testing framework.

  2. Map the hosted provider

    Complete the map-hosted-provider exercise. Separate model creator from billing platform per candidate.

    Open /learn/exercises/map-hosted-provider →

    Output: Hosted-provider mapping note (creator + host + data gap).

  3. Build the side-by-side comparison

    Open /compare/build with the candidate slugs and copy the URL into your notes.

    Open /compare/build →

    Output: Comparison URL that opens the same view for any teammate.

  4. Run the prompt-testing playbook

    Walk the minimum prompt-testing routine against the candidates in your own harness.

    Open /lab/prompt-testing-basics →

    Output: Per-prompt observations recorded against the acceptance rubric.

  5. Run the structured-extraction prompt set

    Open the structured-extraction prompt set and run it against your real schema in your harness.

    Open /lab/prompts/structured-extraction →

    Output: Per-prompt schema validity record + raw responses.

  6. Fill in the prompt test matrix template

    Paste your per-candidate observations into the prompt-test-matrix template.

    Open /lab/templates/prompt-test-matrix →

    Output: Markdown matrix attached to the brief.

  7. Generate the decision evidence brief

    Open /briefs/build with the candidate slugs and export Markdown.

    Open /briefs/build →

    Output: Markdown brief listing verified fields, data gaps, source trail, freshness.

  8. Write the external test plan

    Complete the plan-external-model-test exercise. Pair the brief with workload-specific tests.

    Open /learn/exercises/plan-external-model-test →

    Output: Written test plan covering prompts, latency, rate limits, cost, compliance, regression.

Required resources

The kit reuses existing surfaces — no parallel UI, no duplicated content. Open each surface in the order the timeline lists.

Final checklist

  • Hosted-provider mapping note captured (creator + host + data gap).
  • Comparison URL saved.
  • Prompt-testing observations recorded per prompt, not as a single score.
  • Schema validity record captured for the structured-extraction set.
  • Decision evidence brief exported in Markdown.
  • External test plan written, with regression cadence named.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

Evidence routes

What this kit does not promise

  • Pick the right model for your integration.
  • Assert latency, throughput, or uptime.
  • Certify the model for any regulatory regime.
  • Replace your own workload-specific testing.

What workflow kits do not promise

  • No model recommendations or rankings. Kits walk the evidence; the reader's team makes the decision.
  • No live pricing quotes. Pricing rows referenced inside the kit are sourced references with retrievedAt dates.
  • No production-readiness guarantee. The kit ends with an external test plan — running those tests is the team's responsibility.
  • No compliance certification, legal advice, or vendor endorsement.
  • No SEO ranking guarantees or automation reliability guarantees.
  • No accounts, no progress tracking, no course-completion certificates. Completion is the Markdown artifacts the kit puts in your hands.