Model evaluation plan
> A blank evaluation plan for a single model + workload pairing. Use one copy per candidate.
Identify
- Workload name:
- Candidate model slug (from the catalogue):
- Snapshot identifier or version you intend to test:
- Hosting platform (creator-direct or hosted):
Scope
- Primary use case (link to /use-cases/<slug>):
- Verified fields you intend to weight:
- Data gaps you accept going into the test:
- Out-of-scope behaviours (write them down so they do not creep back in):
Test plan
- Prompt set source and rationale:
- Acceptance rubric per prompt category:
- Sampling parameters held constant (temperature, top_p, max_tokens):
- Region used for latency observation:
- Workload size assumed in the cost projection:
Observations
- Pass / fail summary per category:
- Latency observation (median + tail):
- Cost projection at expected daily volume:
- Notable failure modes:
Decision
- Does the evidence support proceeding to integration, re-scoping, or rejecting? (Write the reasoning, not just the verdict.)
- Open data gaps to close before launch:
- Reviewer sign-offs required:
- Re-test cadence after launch: