# Evaluation plan

## Decision to be made

- Gate concerned:
- System / version / configuration:
- Use case:
- AI task pattern(s):
- Work mode: copilot / bounded automation / strong automation
- Architecture: one model or assistant / tool-assisted workflow / one business agent / orchestrated agent team
- Exact action boundary: A0 / A1 / A2 / A3 / A4
- Interaction / knowledge / deployment profile:
- Jurisdictions and legal triggers:
- Decision-maker:
- Deadline:

## Test set

- Provenance and authorization:
- Development / decision set size:
- Critical segments:
- Difficult, adversarial, and abstention cases:
- Contamination risk:

## Preregistered metrics

| Metric | Baseline | Acceptance threshold | Stop threshold | Segments | Evaluator |
|---|---:|---:|---:|---|---|
| Business outcome | | | | | |
| Accuracy | | | | | |
| Severe error | | | | | |
| Human correction | | | | | |
| Latency | | | | | |
| Cost per outcome | | | | | |

## Pattern-specific profile

Complete every row whose pattern is present. Use `not applicable` only with a
written rationale.

| Pattern | Required measures and tests | Threshold / stop rule |
|---|---|---|
| Generation | accepted quality, factual or source fidelity, prohibited content, reproducibility limits | |
| Retrieval | retrieval coverage, source ACL, groundedness, citation validity, corpus freshness, poisoning | |
| Extraction / classification | confusion matrix, critical classes, abstention, prevalence, subgroup results, drift | |
| Prediction / recommendation | calibration, threshold utility, false-positive and false-negative costs, subgroup results, feedback loops | |
| Conversation | AI disclosure, task completion, human handoff, multi-turn consistency, retention and deletion, abuse | |
| Multimodal | consent and rights, modality rubric, provenance and labelling, transformation robustness, accessibility | |
| Agentic action | plan and tool correctness, authorization, effect read-back, idempotency, rollback, stopping, hostile memory or tool input | |

## Methods

- Deterministic tests:
- Human judgment and rubric:
- Model evaluator and calibration:
- Tool, permission, and effect tests:
- Switzerland transparency, automated-decision, and DPIA checks:
- EU Article 50, role, prohibited-practice, and high-risk checks:
- Treatment of technical incidents:

## Decision

- Result by segment:
- Gaps and uncertainties:
- Accepted / conditionally accepted / rejected:
- Reproducibility and artifacts:

## Consistency and monitoring after changes

- Demonstration cases or real observations:
- Distinct cases / attempts per case / maximum budget:
- Successes per attempt / cases successful on every attempt:
- Failures and retries included in time and cost:
- Sample scored by a human and the automated evaluator / disagreements:
- Reference set to rerun after changes to the model, instructions, tools or data:
- Monitoring owner / frequency / alert threshold / action and fallback:

One success in five attempts is not five successes in five. See the [three validation questions](../docs/evaluations-and-gates.md#three-questions-for-a-reliable-result).
