Templates
Evaluation plan
Decision to be made
- Gate concerned:
- System / version / configuration:
- Use case:
- AI task pattern(s):
- Work mode: copilot / bounded automation / strong automation
- Architecture: one model or assistant / tool-assisted workflow / one business agent / orchestrated agent team
- Exact action boundary: A0 / A1 / A2 / A3 / A4
- Interaction / knowledge / deployment profile:
- Jurisdictions and legal triggers:
- Decision-maker:
- Deadline:
Test set
- Provenance and authorization:
- Development / decision set size:
- Critical segments:
- Difficult, adversarial, and abstention cases:
- Contamination risk:
Preregistered metrics
| Metric | Baseline | Acceptance threshold | Stop threshold | Segments | Evaluator |
|---|---|---|---|---|---|
| Business outcome | |||||
| Accuracy | |||||
| Severe error | |||||
| Human correction | |||||
| Latency | |||||
| Cost per outcome |
Pattern-specific profile
Complete every row whose pattern is present. Use not applicable only with a
written rationale.
| Pattern | Required measures and tests | Threshold / stop rule |
|---|---|---|
| Generation | accepted quality, factual or source fidelity, prohibited content, reproducibility limits | |
| Retrieval | retrieval coverage, source ACL, groundedness, citation validity, corpus freshness, poisoning | |
| Extraction / classification | confusion matrix, critical classes, abstention, prevalence, subgroup results, drift | |
| Prediction / recommendation | calibration, threshold utility, false-positive and false-negative costs, subgroup results, feedback loops | |
| Conversation | AI disclosure, task completion, human handoff, multi-turn consistency, retention and deletion, abuse | |
| Multimodal | consent and rights, modality rubric, provenance and labelling, transformation robustness, accessibility | |
| Agentic action | plan and tool correctness, authorization, effect read-back, idempotency, rollback, stopping, hostile memory or tool input |
Methods
- Deterministic tests:
- Human judgment and rubric:
- Model evaluator and calibration:
- Tool, permission, and effect tests:
- Switzerland transparency, automated-decision, and DPIA checks:
- EU Article 50, role, prohibited-practice, and high-risk checks:
- Treatment of technical incidents:
Decision
- Result by segment:
- Gaps and uncertainties:
- Accepted / conditionally accepted / rejected:
- Reproducibility and artifacts:
Consistency and monitoring after changes
- Demonstration cases or real observations:
- Distinct cases / attempts per case / maximum budget:
- Successes per attempt / cases successful on every attempt:
- Failures and retries included in time and cost:
- Sample scored by a human and the automated evaluator / disagreements:
- Reference set to rerun after changes to the model, instructions, tools or data:
- Monitoring owner / frequency / alert threshold / action and fallback:
One success in five attempts is not five successes in five. See the three validation questions.
To print or save as PDF: Ctrl+P (⌘P on Mac).
Source and history · GitHub