Read-only field-procedure assistant
Helvetia Facilities · internal
Retrieves current authorized passages and drafts a cited answer. No ticket, email, work order, or equipment access.
- 80 frozen questions
- 20 access-boundary tests
- 0 write tools
FIELD GUIDE · AUGUST 2026
Choose one useful problem, prove the value, control the risk, and increase autonomy only when the evidence supports it.
5organization paths
8ordered steps
3safety checks
WHAT THIS PLAYBOOK HELPS YOU DECIDE
The guide separates three questions that are often confused: the job given to the AI, what it may do without you, and the rules that apply where it operates. It then turns those choices into a small, measurable first test.
GUIDED START · ABOUT 3 MINUTES
Answer four short questions. The guide will explain each idea before showing the technical label. Nothing entered here is sent anywhere.
START WITH YOUR REALITY
Step 1/5An independent professional can decide and correct alone. A public service must involve more roles, formal authority, accessibility, and recourse. The useful first pilot is therefore different.
This changes ownership, timing, and the first safeguards. It does not change the core method.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Separate generation, retrieval, prediction, conversation, multimodal work, and action.
PRACTICAL ANSWERS
Each guide gives a direct answer, a comparison, a realistic example, and the sources that limit the claim.
A practical comparison of AI copilots, bounded business agents, and orchestrated agencies, with realistic measures and decision criteria.
02REALISTIC VALUEWhat ROI can a business AI agent realistically deliver?A realistic way to estimate the low and high return of a business AI agent without confusing task speed, workflow automation, and company-wide savings.
03FROM DEMO TO EVIDENCEHow do you pilot a business AI agent?A practical pilot protocol for a bounded business AI agent, from baseline and frozen cases to live evidence, stop rules, and production gates.
04CONTROL BEFORE AUTONOMYHow should a business AI agent be governed?A concise governance model for business AI agents covering ownership, permissions, human approval, evidence, incidents, reassessment, and retirement.
05WHEN ONE AGENT IS NOT ENOUGHWhen does an orchestrated AI agency make sense?A realistic guide to deciding when specialist AI agents and an orchestrator outperform a simpler business agent, with evidence and limits.
06REALISTIC WORKED EXAMPLEWhat does a realistic business AI agent look like for an SME?A worked example of a bounded AI agent preparing B2B quotes for an SME, with eligibility, approval, realistic gains, and failure conditions.
FIRST AXIS · WHAT THE AI ACTUALLY DOES
A chatbot, a predictor, a retrieval system, and an agent can use similar models but require different evidence. Select the dominant pattern, then record every secondary pattern in the use-case card.
FOUR SYNTHETIC CASES · ZERO AUTONOMOUS ACTIONS
RAG, prediction, an external chatbot, and a multimodal assistant can all remain at A0 or A1. Their data, failure modes, legal triggers, and acceptance metrics are still fundamentally different.
Helvetia Facilities · internal
Retrieves current authorized passages and drafts a cited answer. No ticket, email, work order, or equipment access.
Léman Pièces · internal batch
Produces forecasts and intervals for a planner. New products use a manual rule; no supplier order can be created.
Alpina Outdoor · CH + EU
Answers approved public questions, identifies itself as AI, and hands off. No account, payment, refund, or warranty access.
Asteria Home · CH + EU
Reads authorized images and packaging, drafts alt text, and flags mismatches. It cannot edit an asset or publish.
NAME THE SYSTEM BEFORE QUOTING THE GAIN
They move different amounts of work, require different permissions, and must be measured with different outcomes. A percentage without its level is misleading.
MEASUREMENT KEY
Elapsed time from request to result.
Minutes actually spent by a person.
Outputs accepted per owner-hour.
Eligible cases completed without intervention.
A real downstream result, not model activity.
WHAT THE DENOMINATOR CHANGES
Each figure answers a different question. Read the measured outcome and the transfer limit together.
Average issues resolved per hour. Less-experienced workers gained most; top performers had small quality declines.
SMALL BUSINESS · RCT0%About +15% for high performers and −8% for low performers. Access to advice did not guarantee execution quality.
BUSINESS AGENT · RCT5.8%Eligible chats were 16.8% faster, but the whole flow improved 3.2% and customer rating fell on eligible chats.
AUTONOMOUS CODE · FIELD+180%The signal attenuated to +30% releases and no detected increase in total app usage.
AGENCY · BENCHMARK15.2%A 3.5x relative gain over 4.3% at 46 concurrent tasks, in a simulated six-hour environment.
PUBLIC SECTOR · REVIEW19–26 minLarge trials, but no random allocation and no proof that saved time became a delivered public outcome.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Compare one task with measured evidence, then expose every minute that remains human.
MEASURE THE TASK, NOT THE HYPE
Define one repeatable task, inspect a comparable source, and account for preparation, supervision, verification, corrections, exceptions, and setup. External evidence frames a test. Your pilot supplies the answer.
Choose the output unit you can count repeatedly. The organization is deliberately absent because it changes controls, not the task benchmark.
Information search and synthesisFind, compare, and summarize existing information with source checking. one verified answer or synthesis
A comparable source can frame a test. Context-only evidence stays visible but cannot generate a transferable percentage.
Never present the model estimate as measured productivity or use it automatically in the transfer calculation.
Machine runtime is separate. Enter only minutes spent by people, including review and failed cases.
FROM SCENARIO TO PROTOCOL
The calibrator exposes the assumptions. This protocol now fixes the order, practical sample, evidence, and decision before the system touches live work.
Minimum horizon30days
Frozen evaluation set40cases
Bounded live cases20cases
Live collection at this volume≈ 3.1weeks
Write the owner, workflow boundary, baseline, eligibility rule, allowed effects, thresholds, and stop authority.
Run frozen real cases, critical segments, abstentions, adversarial inputs, tool failures, and duplicate events before live use.
Observe the complete workflow with no external effect. Compare accepted outcomes, not model activity, against the manual baseline.
Release only the selected level. Keep human approval, guardian veto, least privilege, logging, and rollback wherever the level requires them.
Judge value, quality, safety, and eligibility separately. Continue the same scope, rework and rerun, or stop and roll back.
ONE DATE · THREE POSSIBLE DECISIONS
All critical gates pass. Keep the same workflow and permissions; set the next review.
Value exists but quality, eligibility, or reliability misses. Fix the cause without increasing autonomy.
A critical gate fails or no useful value appears. Return to the safe process and preserve the evidence.
FROM PILOT TO GATE DECISION
A strong average cannot cancel a critical incident, and missing traces are not a negative result: they make the pilot non-evaluable. The output authorizes one next action, never an automatic increase in autonomy.
OBSERVED PILOT RESULTS
OPERATE WITHOUT LOSING THE BOUNDARY
Translate the gate decision into named ownership, monitoring windows, hard suspension triggers, a rehearsable rollback, and a dated reassessment. A model, tool, permission, policy, or data change reopens the evidence question.
Formal review cadence14days · 1 September 2026
Target time to contain60 minplanning target, rehearse it
Authorized scope1Same proven workflow only
THE FOUR WINDOWS TO WATCH
Every accepted output, correction, rejection, abstention, and exception by segment.
Every tool call, approval, destination, external effect, read-back, duplicate, and rollback result.
Model, prompt, retrieval source, policy, permission, supplier, data mix, latency, and cost changes.
Observed eligibility, human active time, throughput, queue, rework, displaced bottlenecks, and shipped outcome.
SUSPEND IMMEDIATELY WHEN
Any critical, unauthorized, irreversible, misdirected, or untraceable effect occurs.
Required approval, guardian veto, identity boundary, write limit, or fallback is unavailable.
The operating version differs from the evaluated model, prompt, tools, sources, permissions, or policy.
Accepted quality falls below its gate in two consecutive windows, or one protected segment crosses a critical floor.
Cost, latency, queue, or human workload exceeds the written operational limit.
ROLLBACK IN FIVE PROVABLE STEPS
Stop intake and revoke or disable write-capable execution.
Send pending and new cases to the tested manual fallback.
Freeze logs, versions, approvals, tool receipts, destinations, and timestamps.
Read back external state, identify every effect, repair what is safely reversible, and escalate the rest.
Resume only after the owner records cause, corrective action, rerun evidence, and a new gate decision.
Configuration changes are new evidence claims. Cosmetic changes may use a regression check; model, data, retrieval, tool, permission, policy, or workflow changes require the affected frozen tests and gate to be rerun before release.
At the review date, compare against the current manual baseline rather than the original demo. Continue, narrow, replace, or retire. Preserve export, deletion, supplier exit, access revocation, and the manual process.
HAND OFF THE DECISION · NOT THE DEMO
A future owner should be able to reconstruct the assumptions, protocol, observed result, authorized scope, and rollback without relying on memory or a slide deck.
2items still missing
Volume, manual baseline, eligible share, planning range, and setup assumption.
Level, horizon, frozen set, bounded live sample, thresholds, and possible decisions.
Observed value, quality, safety, trace, eligibility, denominator, and authorized next action.
Named owners, scope, monitoring, suspension triggers, rollback, change rule, and review date.
ATTACH OR REFERENCE THESE SIX RECORDS
Signed mandate, scope, affected people, and current manual baseline with denominator.
Exact system inventory: model, prompts, retrieval sources, tools, permissions, policies, suppliers, and versions.
Frozen evaluation-set identifier or hash, segments, adversarial cases, thresholds, and reproducible results.
Live-case ledger with eligibility, approvals, corrections, tool calls, destinations, external effects, read-backs, and rollbacks.
Signed gate decision separating value, quality, safety, evaluability, economics, and authorized scope.
Named operating and incident owners, contact route, fallback proof, rollback rehearsal, next review, and retirement path.
FROM METHOD TO A REAL PILOT
Frame the observation, preserve the full denominator, review the redaction, and export a local draft. Nothing entered here is sent to a server or admitted to the public registry.
One workflow, version, baseline, population, and stop rule.
Keep accepted, failed, excluded, escalated, and missing cases.
An independent person checks provenance, redaction, and transfer limits.
Only reviewed, anonymized reports may enter the public registry.
PRIVACY BOUNDARYLocal-only drafting · Do not enter client names, personal data, secrets, privileged content, raw prompts, or exploitable security details.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Adapt ownership, pace, and safeguards for an independent, company, nonprofit, or public service.
START WITH YOUR REALITY
Same method. Different depth of control, evidence, and responsibility.
YOUR STARTING PLAN · 01
One measured, low-risk workflow with a manual fallback.
Measure five repetitive tasks and exclude high-impact decisions.
Choose the simplest tool and build 20–50 representative tests.
Produce results without sending, publishing, or modifying anything.
Continue, correct, or stop against the written threshold.
Do not skip
ADD SECTOR-SPECIFIC STOP CONDITIONS
Choose the organization path first, then add every sector overlay that touches the service. A hospital can require healthcare and critical-infrastructure gates at the same time.
Owner, baseline, risk, evaluations, pilot.
Name the harm that efficiency cannot offset.
Keep only the authority proven safe in context.
These are operational overlays, not legal classifications. Verify the exact role, jurisdiction, product, population, and sector rules before release.
THE OPERATING LOOP
These are method steps, not release numbers. Open only the step you are working on. Every step ends with concrete evidence, not a presentation.
INTERACTIVE LIFECYCLE
Your entries can be saved in this browser and resumed later. Use non-identifying working information only. The workbench guides a decision; it does not certify compliance.
OrganizationIndependent
Use patternRetrieval
IntegrationIt executes a bounded process
RouteSwitzerland + EU
What observable problem is worth changing, and who may decide?
A tool request without an owner, boundary, and decision date cannot become an accountable project.
Signed mandate with owner, affected people, limits, and decision date.
CONNECTED PROJECT RECORDS
The guide pre-fills linked fields. Edit only what needs a project decision, then keep owners, dates, and evidence references beside the work.
Use patternRetrieval
Legal routeSwitzerland + EU
Keep one operational identity for the system, its purpose, boundaries, owner, supplier, and review route.
A register makes scope and ownership findable without reopening every project discussion.
CHANGE REVIEW
Load an earlier export of this dossier. The comparison stays in this browser, shows one difference at a time, and records the response beside the change.
Export the dossier before a material change, then use that file here as the reference version.
Local project dossier
Answers and selected controls are saved only in this browser. Export the JSON file to move or back up the working dossier.
Do not enter raw client evidence, secrets, or identifying personal data. Browser storage is not an authorized evidence repository.
View the JSON Schema ↗COMPARABLE EXAMPLES
Each case has a different organization, task, autonomy boundary, and proof contract. Select the closest comparison, not the biggest number.
A shared inbox with human review and no automatic send.
WORKED EXAMPLE · FICTIONAL MICRO-BUSINESS
Follow one bounded use case from its four-week baseline to a conditional gate decision. The numbers are synthetic; the evidence structure is reusable.
Atelier Horizon receives quotes, breakdowns, billing questions, and complaints in one shared inbox. The goal is deliberately narrow: suggest routing and prepare a draft. The system never sends or updates anything.
360requests / month
11 minbaseline handling
8 min 35pilot handling
0automatic sends
Time, same-day replies, rework, and routing errors are recorded before choosing a tool.
No automatic send, price promise, CRM write, schedule change, or reply to an ambiguous request.
Forty frozen cases must pass routing, extraction, escalation, unsupported-claim, correction, and time thresholds.
The copilot produces proposals without influencing live replies; every configuration version is recorded.
Three trained users accept, correct, or reject every category and draft before sending.
Value and reliability pass. Automatic sending and system writes remain prohibited while weak segments receive more tests.
GATE 03 · DECISION
It is: keep the measured copilot for 60 more days, review errors weekly, rerun the frozen set after every change, and consider automation only for a stable and reversible subset.
WORKED EXAMPLE 02 · SME · B2B QUOTES · A2
This fictional 42-person industrial SME tests a business agent on one catalogue-quote workflow. The low and high bounds stay visible, excluded requests remain in the denominator, and every price still requires approval.
Noroît Mécanique SA receives quote requests by email with PDFs and spreadsheets. The agent qualifies the request, checks the authorized customer, catalogue, pricing matrix, and lead time, then prepares and verifies the quote. It waits for explicit approval before writing to the ERP and CRM and sending to the displayed recipient.
Known customer, catalogue product, complete units
References, quantities, recipient, requested date
CRM, catalogue, discount matrix, ERP lead time
Deterministic price and margin rules
Facts, conflicts, policy, intended effects
One person sees price, sources, and destination
ERP quote, CRM log, email, and read-back
LOW / CENTRAL / HIGH
160 requests × eligible share × 76 baseline minutes × reduction on eligible work. Capacity is not revenue.
76 → 27 minmedian human time on accepted eligible quotes
×2.81theoretical accepted throughput per human hour
163/238ready to approve without correction
≈ −45%portfolio ceiling after the denominator
0%autonomous completion at A2
OECD: 31% report GenAI use, but only 29% of users report use in core activities. The survey does not measure the size of the gain.
↗EMPIRICAL COPILOT+15%QJE: average increase in resolved support chats per hour across 5,172 workers. Useful lower anchor; not an A2 quote agent.
↗PROVIDER CASE−80 to −95%AWS/Grupo Elfa reports these quote-processing reductions. Useful high anchor; large-scale customer claim, not independent SME proof.
↗These sources make the envelope plausible; they do not validate Noroît’s synthetic result. The local frozen set, live ledger, errors, approvals, full cost, and downstream outcome decide the gate.
GATE 04 · LEVEL DECISION
The observed case is near the central range: 89.8 hours of monthly capacity, about CHF 4,500 net of recurring cost, and a simple setup payback near 3.5 months. Custom parts, exceptional discounts, contracts, and every final price remain human.
WORKED EXAMPLE 03 · FOUNDATION · GRANT DOSSIERS · A2
This fictional 14-person foundation tests an A2 agent on grant administration, never on grant judgment. Every excluded channel stays open, every funding decision stays human, and mission harm overrides productivity.
Fondation Lien Local handles 720 micro-grant applications per year in three languages. The agent checks workflow entry, inventories and cites documents, applies a published completeness checklist, and prepares a pseudonymized reviewer packet. After approval, it writes and routes the packet. It never scores merit, need, or funding probability.
Consent, known program, channel, readable files
Documents and necessary data only
Administrative facts with page citations
Deterministic completeness checklist
Document request or pseudonymized packet
Sources, transformations, and recipients
Grant system, two reviewers, effect read-back
THE A2 AGENT CARRIES
PEOPLE RETAIN
LOW / CENTRAL / HIGH
60 applications × workflow share × 96 baseline minutes × reduction. Setup is CHF 12,000; recurring cost is CHF 750 per month.
96 → 39 minmedian human time per accepted reviewer packet
×2.46theoretical packet throughput per human hour
58/86reviewer-ready without correction
≈ −39%portfolio ceiling after all 120 applications
100%funding decisions made by people
GATE 04 · MISSION BEFORE EFFICIENCY
Telephone, paper, and assisted applications remain available.
The system does not rank vulnerability or infer deservingness.
Rework and stops are reviewed by language, channel, and organization type.
Every decision is explained and can be challenged outside the agent.
Candid: 1% of 529 responding foundations report using GenAI to screen or help decide; 97% say no.
↗FUNCTIONAL ANALOGUE1,000+Degrees of Change handles more than 1,000 applications with 150 volunteer assessors; the provider case describes extraction and staff-reviewed matching, not a causal time result.
↗PROVIDER HIGH BOUND−80%Microsoft reports an 80% reduction in aid-disbursement wait time at NZF. Several changes and a wider automation boundary prevent direct transfer.
↗GATE 05 · AUTONOMY DECISION
The observed case releases 37.5 administrative hours per month, about CHF 1,575 net of recurring cost, with a simple payback near 7.6 months. That creates capacity. It does not mean another grant. Any expansion requires affected-person consultation, larger language and channel samples, and a tested challenge path.
WORKED EXAMPLE · PUBLIC SERVICE · PLANNING DOSSIERS · A2
This fictional Swiss municipality tests an A2 agent on administrative completeness, never on planning judgment. The workflow can move faster only if mandate, procurement, evidence, public notice, appeal, archives, and a no-AI service path move with it.
The City of Mont-Rive receives 1,080 planning applications per year in French and German. The agent inventories documents, extracts source-cited administrative facts, and applies a published checklist. After officer approval, it sends, records, and routes. It never declares completeness, interprets law, or recommends approval, refusal, conditions, or priority.
Authority, signature, channel, known request type
Documents, versions, and necessary data
Parcel and project facts with page or plan citations
Published checklist and controlled official sources
Missing-item list or neutral case packet
Sources, uncertainty, recipient, and effects
Send, register, route, and effect read-back
THE A2 AGENT CARRIES
PEOPLE AND THE AUTHORITY RETAIN
LOW / CENTRAL / HIGH
90 applications × workflow share × 145 baseline minutes × reduction. Setup is CHF 48,000; recurring cost is CHF 3,200 per month, including governance and exit.
145 → 58 minmedian human time per accepted case packet
×2.50theoretical packets per officer-hour
121/166approval-ready without correction
≈ −35%portfolio ceiling across all 270 applications
100%public decisions made by qualified people
P0 → P5
Mandate, baseline, non-AI options, and decision authority.
Applicable law, rights, languages, accessibility, and appeal.
Audit, subprocessors, retention, changes, export, and exit.
Representative cases, segments, security, abuse, and outages.
Shadow mode, named approvers, complaint path, and immediate stop.
Signed decision, public notice, archives, fallback, and withdrawal date.
Milton Keynes officers reported faster validation; receipt-to-validation fell from 15.8 to 7.6 days. The three-month supplier case is not causal proof for a whole service.
↗TASK UPPER BOUND18.5h → 16mCambridge reports 16 minutes for PlanAI summaries versus 18.5 human hours on summaries. Planners still read every submission; this narrow ratio cannot price a dossier workflow.
↗GOVERNANCE ANALOGUE6,000+Leeds handles over 6,000 applications yearly and paired six months of co-design with impact reviews, source links, officer approval, an audit trail, and a public ATRS record.
↗P5 · FORMAL PRODUCTION DECISION
The normalized observation is 76.4 hours of monthly capacity, about CHF 2,757 net of recurring cost, and simple payback near 17.4 months. The low case fails the economic gate. Compliance analysis, reasons, prioritization, or a supplier model change returns the system to P0.
COPILOT CASE · INDEPENDENT PROFESSIONAL · A1
A small pilot should answer a small decision. Follow an independent consultant from meeting notes to a reviewed follow-up, without connecting email, calendar, or client systems.
CAMILLE REY · CLIENT FOLLOW-UP
Camille Rey spends a median 44 minutes turning meeting notes into a summary, action list, and follow-up email. The pilot tests a structured first draft while prices, commitments, recipients, and sending remain exclusively human.
−23%median preparation time
12/14ready within 24 hours
4/14major rework
0invented commitments
Confirm 22 historical follow-ups, the manual fallback, data rules, and an eight-hour setup cap.
Tune on 12 authorized cases, then decide on 12 separate cases against thresholds written in advance.
Generate five drafts but reveal them only after the real follow-up has been written manually.
Review nine live drafts against notes. Add commercial content, choose the recipient, and send manually.
GATE 02 · BOUNDARY DECISION
The median falls from 44 to 34 minutes and all critical gates pass, but 29% of drafts still need major rework. No automatic sending, full proposal generation, or system connection is justified.
WORKED EXAMPLE 03 · BUSINESS AGENT · INDEPENDENT · A2
Phase 2 keeps the same professional, baseline, and outcome. What changes is the system: approved tools, persistent case state, quality control, execution after approval, and explicit exception handling.
After the A1 copilot pilot, Camille tests a business agent on 20 eligible follow-ups. It reads authorized CRM context, prepares the summary and actions, and checks facts and policy. It waits for one explicit approval before updating the CRM, creating tasks, and sending the reviewed email.
Structured notes and eligibility check
Read-only CRM and client rules
Summary, actions, and follow-up
Facts, dates, policy, and conflicts
One informed human decision
Email, CRM, tasks, and audit log
THE AGENT OWNS
THE PERSON OWNS
44 → 14 minmedian human active time · −68%
×3.1accepted follow-ups per owner-hour
13/20ready to approve without correction
3/20correctly escalated
0unapproved external actions
Drafts one step; the person carries and completes the workflow.
Runs the full bounded workflow and executes only after approval.
Autonomous low-risk sending requires 50 more cases and a new gate.
Use separate identity, least privilege, read-only CRM first, idempotent writes, a kill switch, and a tested manual fallback.
Run 40 representative cases, including conflicts, missing context, price requests, prompt injection, duplicate actions, and unavailable tools.
Compare the complete proposed workflow with the real manual follow-up; no email or write reaches a live system.
Camille reviews one evidence packet, approves or refuses, and the agent executes the authorized actions while logging every effect.
GATE 04 · AUTONOMY DECISION
The gain is large because the system now carries the workflow, not because the model merely writes faster. Autonomous sending remains blocked until 50 additional eligible cases show zero critical errors, stable exceptions, no more than 10% major correction, and a verified rollback.
WORKED EXAMPLE 04 · ORCHESTRATED AGENCY · INDEPENDENT · A3
This is where orchestration becomes useful: the work contains distinct research, analysis, quality, and execution roles that can run in parallel. The scope remains one eligible service, not the whole business.
Camille delivers a standardized operational diagnostic for existing small-business clients. After the client interview, the agency qualifies the case, retrieves authorized evidence, scores the process, produces the report and action plan, challenges its own conclusions, then performs low-risk CRM, task, delivery, and scheduling actions inside a pre-approved service policy.
Assigns work, enforces the case policy, resolves dependencies, stops on disagreement, and accepts no specialist’s self-reported success without effect evidence.
Identity, eligibility, minimization
Authorized sources and traceable citations
Diagnosis, scoring, and uncertainty
Report, actions, and client-ready structure
Facts, contradictions, risk, and permissions
Delivery, CRM, tasks, and scheduling
LIKE-FOR-LIKE BENCHMARK
×7.9accepted diagnostics per owner-hour
9/12accepted without major rework
8/12eligible cases completed straight through
4/12stopped and escalated before effect
5h 20median internal cycle vs 18h
0unauthorized commitments or writes
Separate roles, inputs, outputs, permissions, failure boundaries, effect evidence, and the situations that must remain human.
Run 60 cases through manual, copilot, single-agent, and orchestrated conditions; measure accepted output, not agent activity.
Run specialists in parallel without live effects. Inject disagreement, stale memory, tool outages, duplicated events, and poisoned source content.
Allow only catalogued low-risk effects on eligible cases. Guardian veto, cost limits, rollback, and human escalation remain active.
GATE 05 · SCOPE DECISION
The agency may continue for the defined diagnostic and its catalogued effects. It may not choose new services, prices, contracts, clients, data classes, or permissions. Broad multi-system autonomy requires a separate mandate, an independent audit, and evidence across several workflows.
TECHNICAL PROGRESSION
Move one level at a time. Stop when a simpler system meets the need.
QUICK CONTROL ORIENTATION
This is internal triage, not a legal classification.
VERSIONED CONTROL CROSSWALK · JSON 1.1
Organization, impact, autonomy, use pattern, and jurisdiction resolve to stable control IDs. Each row exposes its trigger, evidence, decision gates, lifecycle phases, and dated source references.
18candidate controls
Independentorganization
R1 × A1active profile
Retrievaluse pattern
Switzerland + EUjurisdiction route
12versioned sources
Tie the system to one measurable problem, a bounded mandate, prohibited actions, and a person accountable for the final decision.
Make official and informal AI use visible, with ownership, provider, model, risk, autonomy, and review dates.
Compare the AI system with the real current process using complete denominators and a business or public-service outcome.
Prevent a low technical autonomy label from hiding high human impact, sensitive data, or regulated responsibility.
Record provenance, purpose, access, processing locations, reuse, retention, deletion, and subprocessors before using personal or confidential data.
Trigger conditionWhen personal or confidential data is processed
Preserve audit, change notice, export, deletion, continuity, and exit rights when an external provider is part of the system.
Trigger conditionWhen an external provider is used
Fix the decision unit, test population, thresholds, critical errors, and authorized judge before optimizing or piloting.
Keep a decision set separate from development and reject averages that conceal failure on a critical segment or an unacceptable error.
Trigger conditionWhen people may be materially affected
Use attributable identities, read-only defaults, temporary permissions, approved destinations, isolated secrets, and enforceable limits.
Trigger conditionWhen external content or tools cross a trust boundary
Test hostile content, identity confusion, poisoned context, unsafe tool parameters, partial failure, loops, data leakage, and attempts to disable controls.
Trigger conditionWhen external content or tools cross a trust boundary
Ensure a qualified person has enough information, time, authority, and an operational channel to reject, correct, override, and hear appeals.
Trigger conditionWhen people may be materially affected
Link each case to the exact configuration, evidence, approval, tool parameters, external effect, read-back, actor, and time.
Trigger conditionWhen people may be materially affected
Give a reachable owner the authority and tested procedure to contain, preserve evidence, notify, revoke access, recover, and trigger reassessment.
Trigger conditionWhen external content or tools cross a trust boundary
Monitor business outcomes, critical segments, corrections, incidents, drift, and cost; return to the right gate after any material model, data, tool, permission, or population change.
Stop the system cleanly by revoking access, exporting what continuity requires, disposing of data, preserving required evidence, and restoring a safe process.
Record task, interaction, knowledge, modality, deployment, operating mode, effect, and jurisdiction separately from impact and autonomy before selecting a system.
Add retrieval, classification, prediction, conversation, multimodal, or agentic measures and stop rules whenever those patterns are present.
Extend the threat model for retrieval, predictive models, conversation, multimodal inputs, code execution, persistent memory, and inter-agent communication as applicable.
USE THE PLAYBOOK
Copy the operational templates, complete the first gate, and keep the evidence with the project.
Owner, baseline, outcome, and boundaries.
↗02Value and difficulty kept separate.
↗03Scenarios, controls, and residual risk.
↗04Metrics, segments, thresholds, and stop rules.
↗05Value, reliability, and risk judged separately.
↗06Contain, qualify, recover, and learn.
↗07Complete tasks, assistive technologies, and an equivalent non-AI channel.
↗08Affected groups, rights, safeguards, recourse, and residual impact.
↗09Observed evidence, anonymization review, and explicit transfer limits.
↗