MUSYG · AI ADOPTION

Templates

Evaluation plan

Decision to be made

  • Gate concerned:
  • System / version / configuration:
  • Use case:
  • AI task pattern(s):
  • Work mode: copilot / bounded automation / strong automation
  • Architecture: one model or assistant / tool-assisted workflow / one business agent / orchestrated agent team
  • Exact action boundary: A0 / A1 / A2 / A3 / A4
  • Interaction / knowledge / deployment profile:
  • Jurisdictions and legal triggers:
  • Decision-maker:
  • Deadline:

Test set

  • Provenance and authorization:
  • Development / decision set size:
  • Critical segments:
  • Difficult, adversarial, and abstention cases:
  • Contamination risk:

Preregistered metrics

Metric Baseline Acceptance threshold Stop threshold Segments Evaluator
Business outcome
Accuracy
Severe error
Human correction
Latency
Cost per outcome

Pattern-specific profile

Complete every row whose pattern is present. Use not applicable only with a written rationale.

Pattern Required measures and tests Threshold / stop rule
Generation accepted quality, factual or source fidelity, prohibited content, reproducibility limits
Retrieval retrieval coverage, source ACL, groundedness, citation validity, corpus freshness, poisoning
Extraction / classification confusion matrix, critical classes, abstention, prevalence, subgroup results, drift
Prediction / recommendation calibration, threshold utility, false-positive and false-negative costs, subgroup results, feedback loops
Conversation AI disclosure, task completion, human handoff, multi-turn consistency, retention and deletion, abuse
Multimodal consent and rights, modality rubric, provenance and labelling, transformation robustness, accessibility
Agentic action plan and tool correctness, authorization, effect read-back, idempotency, rollback, stopping, hostile memory or tool input

Methods

  • Deterministic tests:
  • Human judgment and rubric:
  • Model evaluator and calibration:
  • Tool, permission, and effect tests:
  • Switzerland transparency, automated-decision, and DPIA checks:
  • EU Article 50, role, prohibited-practice, and high-risk checks:
  • Treatment of technical incidents:

Decision

  • Result by segment:
  • Gaps and uncertainties:
  • Accepted / conditionally accepted / rejected:
  • Reproducibility and artifacts:

Consistency and monitoring after changes

  • Demonstration cases or real observations:
  • Distinct cases / attempts per case / maximum budget:
  • Successes per attempt / cases successful on every attempt:
  • Failures and retries included in time and cost:
  • Sample scored by a human and the automated evaluator / disagreements:
  • Reference set to rerun after changes to the model, instructions, tools or data:
  • Monitoring owner / frequency / alert threshold / action and fallback:

One success in five attempts is not five successes in five. See the three validation questions.

To print or save as PDF: Ctrl+P (⌘P on Mac).

Source and history · GitHub