Prompt test matrix
> A row-per-prompt matrix you fill in per candidate model. Pair with the model evaluation plan.
Matrix legend
- Use ✓ for pass against the acceptance rubric, ✗ for fail, ? for ambiguous (record why).
- Record latency in seconds wall-clock from your environment.
- Record cost as input + output tokens; convert to currency later using the catalogue's pricing reference.
Prompt index
- | ID | Prompt category | Prompt summary | Acceptance rubric |
- | --- | --- | --- | --- |
- | P-01 | happy path | | |
- | P-02 | edge case | | |
- | P-03 | adversarial | | |
- | P-04 | refusal | | |
- | P-05 | long context | | |
Results per candidate
- | Prompt ID | Candidate A | Candidate B | Notes |
- | --- | --- | --- | --- |
- | P-01 | | | |
- | P-02 | | | |
- | P-03 | | | |
- | P-04 | | | |
- | P-05 | | | |
Rollup
- Pass rate per candidate (count, not percentage):
- Median latency per candidate:
- Notable failure modes per candidate:
- Open data gaps surfaced by this matrix: