Kit · governance teams
Governance review kit
Prepare a source / freshness / lifecycle review package for internal governance discussions. Walks the governance learning path, four lessons (lifecycle, status, benchmark limits, pricing), three exercises that produce the freshness checklist + lifecycle note + test plan, the regression + prompt-testing playbooks, the refusal-boundary + instruction-following prompt sets, and the evaluation plan + automation risk checklist templates.
Audience
governance teams
Difficulty
intermediate
Estimated time
200 min · 9 steps
Who this kit is for
Lifecycle gates, status observations kept separate from vendor claims, source freshness checks, and Markdown evidence trails for risk, compliance, and governance reviewers. The platform never certifies a model — it surfaces the evidence your team owns the verdict on.
Goal
End with a source freshness checklist, a lifecycle review note, an explicit data-gap list, a Markdown governance review brief, and a written external testing plan — never a certification.
What you will produce
- Source freshness checklist
- Lifecycle review note
- Data gap list
- Governance review brief
- External test plan
Prerequisites
Export Markdown
The kit serialises to a single Markdown document you can paste into a design doc, ticket, or PR description.
Workflow
9 sequenced steps
Open each step in order. Every step opens a route that already exists — no parallel UI.
Step-by-step workflow
Walk the governance learning path
Read the four lessons (model-lifecycle, status-aware-selection, benchmark-limitations, pricing-references).
Output: Notes covering lifecycle, status, benchmark limits, pricing references.
Check source freshness
Open the reverification queue filtered to the provider under review; export the checklist.
Open
/learn/exercises/check-source-freshness→Output: Source freshness checklist (Markdown / JSON via the export endpoint).
Inspect model lifecycle
Capture lifecycle state + retirement date + source citation for each model in scope.
Open
/learn/exercises/inspect-model-lifecycle→Output: Lifecycle review note covering each model + the provider's deprecation history.
Walk the coverage audit
Open /coverage filtered to the provider; list every unverified field that needs external follow-up.
Output: Explicit data-gap list for the review board.
Run the model regression testing playbook
Freeze the canary suite that will catch silent snapshot drift after the review approves.
Open
/lab/model-regression-testing→Output: Canary suite + documented regression cadence.
Run the refusal-boundary prompt set
Run the refusal-boundary prompts in your harness to surface over-refusal and inappropriate-compliance behaviour.
Open
/lab/prompts/refusal-boundary→Output: Per-prompt observations recorded for the review record.
Fill the evaluation plan + automation risk checklist templates
Adapt both templates to the review's scope; paste observations from earlier steps.
Open
/lab/templates/model-evaluation-plan→Output: Two filled Markdown templates attached to the review package.
Generate the governance review brief
Open /briefs/build with the models in scope and export Markdown for the review board.
Output: Markdown brief with verified fields + data gaps + freshness + lifecycle.
Write the external test plan
Complete the plan-external-model-test exercise to pair the review with workload-specific tests.
Open
/learn/exercises/plan-external-model-test→Output: Written external test plan to attach to the review package.
Required resources
The kit reuses existing surfaces — no parallel UI, no duplicated content. Open each surface in the order the timeline lists.
Lessons
Model lifecycle: active, deprecated, retired →
What active, preview, deprecated, and retired mean for a model — and why lifecycle should gate integration decisions.
Status-aware model selection →
Why vendor-reported status pages and independent probes are kept separate — and when status should gate a model decision.
Why benchmark scores can mislead →
Contamination, prompt variance, version drift, and why the catalogue does not publish provider-reported benchmark scores casually.
AI model pricing references explained →
Why catalogue pricing rows are references, not quotes — and how to read them without ranking models by price.
Exercises
Check source freshness and reverification state →
Open the sources hub for a provider, identify any stale citations, and walk the reverification queue to see what is due for re-check.
Inspect lifecycle before integration →
Pull lifecycle state for a candidate model, check for retirement date, and add a migration target to your notes if one exists.
Plan an external model test →
Use the brief and the testing lesson to plan your own prompt, latency, rate-limit, and cost validation work for the shortlist.
Lab playbooks
Model regression testing →
How to run a small, repeatable canary suite after every snapshot rotation so silent regressions surface before production traffic notices.
Prompt testing basics →
The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.
Evaluation prompt sets
Refusal boundary →
Evaluate whether the model handles benign boundary-setting safely — without over-refusing, without giving definitive professional advice, and without complying with inappropriate requests.
Instruction following →
Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.
Final checklist
- Source freshness checklist exported per provider in scope.
- Lifecycle review note covers every model in scope.
- Data gap list is explicit (not implied).
- Canary suite + regression cadence are documented.
- Governance review brief exported in Markdown.
- External test plan written.
- No certification or compliance-approval language in any artifact.
Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.
Evidence routes
What this kit does not promise
- Compliance approval or certification.
- Legal advice.
- Vendor endorsement.
- Sign-off on behalf of the reviewer's organisation.
What workflow kits do not promise
- No model recommendations or rankings. Kits walk the evidence; the reader's team makes the decision.
- No live pricing quotes. Pricing rows referenced inside the kit are sourced references with retrievedAt dates.
- No production-readiness guarantee. The kit ends with an external test plan — running those tests is the team's responsibility.
- No compliance certification, legal advice, or vendor endorsement.
- No SEO ranking guarantees or automation reliability guarantees.
- No accounts, no progress tracking, no course-completion certificates. Completion is the Markdown artifacts the kit puts in your hands.