Skip to content
WebmasterID

Lab · Template

Prompt test matrix

A row-per-prompt matrix you fill in per candidate model. Pair with the model evaluation plan.

Export

Open or pipe the raw Markdown into your design doc, ticket, or PR description.

Open raw Markdown → /api/lab/templates/prompt-test-matrix

Template

4 sections

Render order matches the Markdown export. Each section is generic — adapt to your workload before filling in.

Prompt test matrix

> A row-per-prompt matrix you fill in per candidate model. Pair with the model evaluation plan.

Matrix legend

  • Use ✓ for pass against the acceptance rubric, ✗ for fail, ? for ambiguous (record why).
  • Record latency in seconds wall-clock from your environment.
  • Record cost as input + output tokens; convert to currency later using the catalogue's pricing reference.

Prompt index

  • | ID | Prompt category | Prompt summary | Acceptance rubric |
  • | --- | --- | --- | --- |
  • | P-01 | happy path | | |
  • | P-02 | edge case | | |
  • | P-03 | adversarial | | |
  • | P-04 | refusal | | |
  • | P-05 | long context | | |

Results per candidate

  • | Prompt ID | Candidate A | Candidate B | Notes |
  • | --- | --- | --- | --- |
  • | P-01 | | | |
  • | P-02 | | | |
  • | P-03 | | | |
  • | P-04 | | | |
  • | P-05 | | | |

Rollup

  • Pass rate per candidate (count, not percentage):
  • Median latency per candidate:
  • Notable failure modes per candidate:
  • Open data gaps surfaced by this matrix:

Policy: The matrix does not score candidates against each other. Fill the rollup with observations the reviewer can read; do not collapse the evidence into a single number.

Generated by WebmasterID Models AI Usage Lab. No fabricated metrics. No model recommendations. /lab/templates

Related playbooks

Pair with evaluation prompt sets

The matrix is generic by design. Fill the row IDs from an evaluation prompt set so the column labels stay traceable.

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.