Read-only field-procedure assistant
Helvetia Facilities · internal
Retrieves current authorized passages and drafts a cited answer. No ticket, email, work order, or equipment access.
- 80 frozen questions
- 20 access-boundary tests
- 0 write tools
FIELD GUIDE · SEPTEMBER 2026
Choose one useful problem, prove the value, control the risk, and increase autonomy only when the evidence supports it.
5organization starting paths
8adoption method steps
3checks before moving on
WHAT THIS PLAYBOOK HELPS YOU DECIDE
The guide separates ideas that are often confused: the kind of AI task, how people and AI share the work, what the system is made of, and what it may do without you. It then checks the rules for the territory and turns the route into a small, measurable first test.
GUIDED START · ABOUT 3 MINUTES
Answer four questions, then receive your starting plan on the fifth screen. The operating-design screen contains three separate choices. The guide explains each idea before showing a technical code. Nothing entered here is sent anywhere.
START WITH YOUR REALITY
Step 1/5An independent professional can decide and correct alone. A public service must involve more roles, formal authority, accessibility, and recourse. The useful first pilot is therefore different.
This changes ownership, timing, and the first safeguards. It does not change the core method.
The five opening screens choose your starting point. The pilot workspace contains six topics, ending with optional feedback. The eight method steps explain the overall adoption process; they are a reference, not eight more screens to complete.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Separate generation, retrieval, prediction, conversation, multimodal work, and action.
PRACTICAL ANSWERS
Each guide gives a direct answer, a comparison, a realistic example, and the sources that limit the claim.
A practical comparison of AI copilots, bounded business agents, and orchestrated agencies, with realistic measures and decision criteria.
02REALISTIC VALUEWhat ROI can a business AI agent realistically deliver?A realistic way to estimate the low and high return of a business AI agent without confusing task speed, workflow automation, and company-wide savings.
03FROM DEMO TO EVIDENCEHow do you pilot a business AI agent?A practical pilot protocol for a bounded business AI agent, from baseline and frozen cases to live evidence, stop rules, and production gates.
04CONTROL BEFORE AUTONOMYHow should a business AI agent be governed?A concise governance model for business AI agents covering ownership, permissions, human approval, evidence, incidents, reassessment, and retirement.
05WHEN ONE AGENT IS NOT ENOUGHWhen does an orchestrated AI agency make sense?A realistic guide to deciding when specialist AI agents and an orchestrator outperform a simpler business agent, with evidence and limits.
06REALISTIC WORKED EXAMPLEWhat does a realistic business AI agent look like for an SME?A worked example of a bounded AI agent preparing B2B quotes for an SME, with eligibility, approval, realistic gains, and failure conditions.
FIRST AXIS · WHAT THE AI ACTUALLY DOES
A chatbot, a predictor, a retrieval system, and an agent can use similar models but require different evidence. Select the dominant pattern, then record every secondary pattern in the use-case card.
FOUR SYNTHETIC CASES · ZERO AUTONOMOUS ACTIONS
RAG, prediction, an external chatbot, and a multimodal assistant can all remain at A0 or A1. Their data, failure modes, legal triggers, and acceptance metrics are still fundamentally different.
Helvetia Facilities · internal
Retrieves current authorized passages and drafts a cited answer. No ticket, email, work order, or equipment access.
Léman Pièces · internal batch
Produces forecasts and intervals for a planner. New products use a manual rule; no supplier order can be created.
Alpina Outdoor · CH + EU
Answers approved public questions, identifies itself as AI, and hands off. No account, payment, refund, or warranty access.
Asteria Home · CH + EU
Reads authorized images and packaging, drafts alt text, and flags mismatches. It cannot edit an asset or publish.
NAME THE SYSTEM BEFORE QUOTING THE GAIN
First ask how people and AI share the work. Then describe what the system is made of. Finally, set its exact right to act. A0 to A4 describe authority, not the number of agents. The same architecture can operate at different autonomy levels.
Who does each part? In copilot mode, the person drives every step. With bounded automation, the system completes a defined part. With strong automation, it completes most eligible work while people set limits and handle exceptions.
What is connected? It may be one model, fixed tool-assisted steps, one agent choosing steps, or several specialist agents coordinated together.
How far may it act? A0 only advises. A1 searches or drafts. A2 acts after explicit approval. A3 performs bounded actions alone. A4 has broad multi-system authority.
What could go wrong? R0 is a low-impact internal use. R1 is reviewed assistance. R2 involves people, personal data, or an external effect. R3 can affect rights, health, employment, safety, or public authority.
A published result can inform your estimate when the task, finish level, human experience, and way of using AI are close enough. It remains a starting range, not a promise.
MEASUREMENT KEY
Elapsed time from request to result.
Minutes actually spent by a person.
Outputs accepted per owner-hour.
Eligible cases completed without intervention.
A real downstream result, not model activity.
WHAT THE DENOMINATOR CHANGES
Each figure answers a different question. Read the measured outcome and the transfer limit together.
Average issues resolved per hour. Less-experienced workers gained most; top performers had small quality declines.
SMALL BUSINESS · RCT0%About +15% for high performers and −8% for low performers. Access to advice did not guarantee execution quality.
BUSINESS AGENT · RCT5.8%Eligible chats were 16.8% faster, but the whole flow improved 3.2% and customer rating fell on eligible chats.
AUTONOMOUS CODE · FIELD+180%The signal attenuated to +30% releases and no detected increase in total app usage.
AGENCY · BENCHMARK15.2%A 3.5x relative gain over 4.3% at 46 concurrent tasks, in a simulated six-hour environment.
PUBLIC SECTOR · REVIEW19–26 minLarge trials, but no random allocation and no proof that saved time became a delivered public outcome.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Compare one task with measured evidence, then expose every minute that remains human.
MEASURE THE TASK, NOT THE HYPE
Define one repeatable task, inspect a comparable source, and account for preparation, supervision, verification, corrections, exceptions, and setup. External evidence frames a test. Your pilot supplies the answer.
Choose one precise, repeated task, such as drafting an email or reviewing a case. We compare the work first, not the size of the organization.
Information search and synthesisFind, compare, and summarize existing information with source checking. one verified answer or synthesis
We look for research on work close to yours. A sufficiently similar study can suggest a starting range. Other studies remain useful examples but do not change the calculation.
A means the strongest direct comparison in this register. E means a synthetic estimate or planning assumption. The selected source is highlighted below.
Observed comparison between work with and without AI
Operational measure from real work
Time reported by users
Published case or capability test
Synthetic estimate or planning assumption
The figure remains visible for context, but it is not added to your estimated saving.
Machine runtime is separate. Enter only minutes spent by people, including review and failed cases.
The main fields describe the middle case. Add work for the cautious case and remove work for the favourable case. Defaults are examples, not study findings. Review and setup cannot fall below zero; exception rates stay between 0% and 100%.
Add only work missing from both the study and your breakdown. If study time includes setup, enter that part per transferred case: it is removed before adding your local setup. Leave zero when setup is excluded; check the source when unsure.
The work mode changes the setup effort and which studies are comparable. It does not add a fixed productivity bonus. Architecture and A0 to A4 authority are chosen separately. For each case, the current assumptions allow this much remaining human work: 33 min. Before counting setup, this represents 45% less human time. The net result then adds 7.1 min per case during the period you chose for spreading the setup effort.
Start with clearly labelled demonstration cases if needed. Set a budget before testing. In live use, assign someone to review alerts and update the tests.
FROM SCENARIO TO PROTOCOL
Write down what will count as success or failure before testing. Choose the cases to compare, who checks the results, and when to stop. The steps below produce a test brief you can save.
The numbers below are starting minimums used by this guide, not a universal sample size. Use more cases when failures are rare or consequences matter more.
Starting duration30days
Same cases for each version40cases
Real cases within the agreed scope20cases
Live collection at this volume≈ 3.1weeks
Name the task, its current time, which cases the test covers, what counts as success, and who can stop it. Example: draft a reply without sending it.
Use the same set of cases for each version. Include normal requests, missing data, deliberate misleading inputs, tool failures, and duplicate requests. Do not change real systems yet.
Observe the complete workflow with no external effect. Compare accepted outcomes, not model activity, against the manual baseline.
Release only the selected level. Keep human approval, guardian veto, least privilege, logging, and rollback wherever the level requires them.
Judge value, quality, safety, and eligibility separately. Continue the same scope, rework and rerun, or stop and roll back.
ONE DATE · THREE POSSIBLE DECISIONS
All critical gates pass. Keep the same workflow and permissions; set the next review.
Value exists but quality, eligibility, or reliability misses. Fix the cause without increasing autonomy.
A critical gate fails or no useful value appears. Return to the safe process and preserve the evidence.
FROM PILOT TO GATE DECISION
A strong average cannot cancel a critical incident, and missing traces are not a negative result: they make the pilot non-evaluable. The output authorizes one next action, never an automatic increase in autonomy.
OBSERVED PILOT RESULTS
This fills a coherent fictional pilot so you can see how the decision logic works. Every value remains editable.
Observed share eligiblen/acalculated from eligible cases ÷ all observed requests
OPERATE WITHOUT LOSING THE BOUNDARY
Name who checks the results, who can stop the system, and when to review it. Test how to resume work manually after a problem. In a small organization, one person may hold several roles if they can actually carry them out. A change to the model, tools, permissions, rules, or data requires another review.
Formal review cadence14days · Not set
Target time to contain1 hplanning target, rehearse it
Authorized scope1Same proven workflow only
THE FOUR WINDOWS TO WATCH
Every accepted output, correction, rejection, abstention, and exception by segment.
Every tool call, approval, destination, external effect, read-back, duplicate, and rollback result.
Model, prompt, retrieval source, policy, permission, supplier, data mix, latency, and cost changes.
Observed eligibility, human active time, throughput, queue, rework, displaced bottlenecks, and shipped outcome.
SUSPEND IMMEDIATELY WHEN
Any critical, unauthorized, irreversible, misdirected, or untraceable effect occurs.
Required approval, guardian veto, identity boundary, write limit, or fallback is unavailable.
The operating version differs from the evaluated model, prompt, tools, sources, permissions, or policy.
Accepted quality falls below its gate in two consecutive windows, or one protected segment crosses a critical floor.
Cost, latency, queue, or human workload exceeds the written operational limit.
ROLLBACK IN FIVE PROVABLE STEPS
Stop intake and revoke or disable write-capable execution.
Send pending and new cases to the tested manual fallback.
Freeze logs, versions, approvals, tool receipts, destinations, and timestamps.
Check what the tool has already changed. Undo only changes that can be safely reversed and send the rest to the responsible person.
Resume only after the owner records cause, corrective action, rerun evidence, and a new gate decision.
Configuration changes are new evidence claims. Cosmetic changes may use a regression check; model, data, retrieval, tool, permission, policy, or workflow changes require the affected frozen tests and gate to be rerun before release.
At the review date, compare against the current manual baseline rather than the original demo. Continue, narrow, replace, or retire. Preserve export, deletion, supplier exit, access revocation, and the manual process.
HAND OFF THE DECISION · NOT THE DEMO
A future owner should be able to reconstruct the assumptions, protocol, observed result, authorized scope, and rollback without relying on memory or a slide deck.
7items still missing
Volume, manual baseline, eligible share, planning range, and setup assumption.
Level, horizon, frozen set, bounded live sample, thresholds, and possible decisions.
Observed value, quality, safety, trace, eligibility, denominator, and authorized next action.
Named owners, scope, monitoring, suspension triggers, rollback, change rule, and review date.
ATTACH OR REFERENCE THESE SIX RECORDS
Signed mandate, scope, affected people, and current manual baseline with denominator.
Exact system inventory: model, prompts, retrieval sources, tools, permissions, policies, suppliers, and versions.
Frozen evaluation-set identifier or hash, segments, adversarial cases, thresholds, and reproducible results.
Live-case ledger with eligibility, approvals, corrections, tool calls, destinations, external effects, read-backs, and rollbacks.
Signed gate decision separating value, quality, safety, evaluability, economics, and authorized scope.
Named operating and incident owners, contact route, fallback proof, rollback rehearsal, next review, and retirement path.
OPTIONAL CONTRIBUTION
Your own pilot and local dossier do not require a contribution to this public project. If you choose to contribute, compare your estimate with observations across all requests, including failures. An estimate remains distinct from a measured result. The cohort count below tracks contributions, not your progress.
Record the source, transfer contract, net range, and local assumptions before observing outcomes.
Keep accepted, failed, excluded, escalated, and missing cases in the denominator.
Show whether the observed whole-workload result falls below, within, or above the planned range.
An independent person checks provenance, redaction, limits, and whether the result may enter the registry.
PRIVACY BOUNDARYLocal-only drafting · Do not enter client names, personal data, secrets, privileged content, raw prompts, or exploitable security details.
ONE TOPIC AT A TIME
Only the selected topic appears below. Your previous choices stay available while you explore.
Adapt ownership, pace, and safeguards for an independent, company, nonprofit, or public service.
START WITH YOUR REALITY
Same method. Different depth of control, evidence, and responsibility.
YOUR STARTING PLAN · 01
One measured, low-risk workflow with a manual fallback.
Measure five repetitive tasks and exclude high-impact decisions.
Choose the simplest tool and build 20–50 representative tests.
Produce results without sending, publishing, or modifying anything.
Continue, correct, or stop against the written threshold.
Do not skip
ADD SECTOR-SPECIFIC STOP CONDITIONS
Choose the organization path first, then add every sector overlay that touches the service. A hospital can require healthcare and critical-infrastructure gates at the same time.
Owner, baseline, risk, evaluations, pilot.
Name the harm that efficiency cannot offset.
Keep only the authority proven safe in context.
These are operational overlays, not legal classifications. Verify the exact role, jurisdiction, product, population, and sector rules before release.
THE OPERATING LOOP
The map shows which lifecycle phases belong to each method step. Open the step to understand its purpose, then complete the corresponding phases in the workspace below.
Name the owner, the observable problem, the people affected, the limits, and the decision date.
Evidence to keepA signed mandate and a measurable baseline.
Observe decisions, exceptions, data, systems, suppliers, waiting time, and informal AI use.
Evidence to keepA current-state map and AI system register.
Score value and difficulty separately. Start with a frequent, measurable, reversible case.
Evidence to keepComparable use-case cards and one chosen pilot.
Assess impact, data, scale, jurisdictions, reversibility, and the powers granted to the system.
Evidence to keepA documented classification and required reviews.
Test rules and conventional automation before retrieval, tool use, agents, or multi-agent designs.
Evidence to keepAn architecture decision and supplier assessment.
Use real cases, critical segments, adversarial inputs, abstentions, and thresholds written before the pilot.
Evidence to keepA reproducible evaluation plan with stop criteria.
Move from shadow mode to human-approved copilot, then bounded automation only after each gate passes.
Evidence to keepA pilot decision separating value, reliability, and risk.
Version everything, monitor outcomes, rehearse incidents, preserve a manual fallback, and plan withdrawal.
Evidence to keepRunbooks, review date, rollback, and retirement plan.
INTERACTIVE LIFECYCLE
Your entries can be saved in this browser and resumed later. Use non-identifying working information only. The workbench guides a decision; it does not certify compliance.
OrganizationIndependent
Use patternRetrieval
Work modeBounded automation
ArchitectureTool-assisted workflow
RouteSwitzerland + EU
What observable problem is worth changing, and who may decide?
A tool request without an owner, boundary, and decision date cannot become an accountable project.
Signed mandate with owner, affected people, limits, and decision date.
CONNECTED PROJECT RECORDS
The guide pre-fills linked fields. Edit only what needs a project decision, then keep owners, dates, and evidence references beside the work.
Use patternRetrieval
Legal routeSwitzerland + EU
Keep one operational identity for the system, its purpose, boundaries, owner, supplier, and review route.
A register makes scope and ownership findable without reopening every project discussion.
CHANGE REVIEW
Load an earlier export of this dossier. The comparison stays in this browser, shows one difference at a time, and records the response beside the change.
Export the dossier before a material change, then use that file here as the reference version.
Local project dossier
Answers and selected controls are saved only in this browser. Export the JSON file to move or back up the working dossier.
Do not enter raw client evidence, secrets, or identifying personal data. Browser storage is not an authorized evidence repository.
View the JSON Schema ↗COMPARABLE EXAMPLES
Each case has a different organization, task, autonomy boundary, and proof contract. Select the closest comparison, not the biggest number.
A shared inbox with human review and no automatic send.
WORKED EXAMPLE 01 · FICTIONAL MICRO-BUSINESS
Follow one bounded use case from its four-week baseline to a conditional gate decision. The numbers are synthetic; the evidence structure is reusable.
Atelier Horizon receives quotes, breakdowns, billing questions, and complaints in one shared inbox. The goal is deliberately narrow: suggest routing and prepare a draft. The system never sends or updates anything.
360requests / month
11 minbaseline handling
8 min 35pilot handling
0automatic sends
Time, same-day replies, rework, and routing errors are recorded before choosing a tool.
No automatic send, price promise, CRM write, schedule change, or reply to an ambiguous request.
Forty frozen cases must pass routing, extraction, escalation, unsupported-claim, correction, and time thresholds.
The copilot produces proposals without influencing live replies; every configuration version is recorded.
Three trained users accept, correct, or reject every category and draft before sending.
Value and reliability pass. Automatic sending and system writes remain prohibited while weak segments receive more tests.
CASE DECISION
It is: keep the measured copilot for 60 more days, review errors weekly, rerun the frozen set after every change, and consider automation only for a stable and reversible subset.
WORKED EXAMPLE 02 · SME · B2B QUOTES · A2
This fictional 42-person industrial SME tests a business agent on one catalogue-quote workflow. The low and high bounds stay visible, excluded requests remain in the denominator, and every price still requires approval.
Noroît Mécanique SA receives quote requests by email with PDFs and spreadsheets. The agent qualifies the request, checks the authorized customer, catalogue, pricing matrix, and lead time, then prepares and verifies the quote. It waits for explicit approval before writing to the ERP and CRM and sending to the displayed recipient.
Known customer, catalogue product, complete units
References, quantities, recipient, requested date
CRM, catalogue, discount matrix, ERP lead time
Deterministic price and margin rules
Facts, conflicts, policy, intended effects
One person sees price, sources, and destination
ERP quote, CRM log, email, and read-back
LOW / CENTRAL / HIGH
160 requests × eligible share × 76 baseline minutes × reduction on eligible work. Capacity is not revenue.
76 → 27 minmedian human time on accepted eligible quotes
×2.81theoretical accepted throughput per human hour
163/238ready to approve without correction
≈ −45%portfolio ceiling after the denominator
0%autonomous completion at A2
OECD: 31% report GenAI use, but only 29% of users report use in core activities. The survey does not measure the size of the gain.
↗EMPIRICAL COPILOT+15%QJE: average increase in resolved support chats per hour across 5,172 workers. Useful lower anchor; not an A2 quote agent.
↗PROVIDER CASE−80 to −95%AWS/Grupo Elfa reports these quote-processing reductions. Useful high anchor; large-scale customer claim, not independent SME proof.
↗These sources make the envelope plausible; they do not validate Noroît’s synthetic result. The local frozen set, live ledger, errors, approvals, full cost, and downstream outcome decide the gate.
CASE AUTONOMY DECISION
In this synthetic scenario, the result is near the central range: 89.8 hours of monthly capacity, about CHF 4,500 net of recurring cost, and a simple setup payback near 3.5 months. Custom parts, exceptional discounts, contracts, and every final price remain human.
WORKED EXAMPLE 03 · FOUNDATION · GRANT DOSSIERS · A2
This fictional 14-person foundation tests an A2 agent on grant administration, never on grant judgment. Every excluded channel stays open, every funding decision stays human, and mission harm overrides productivity.
Fondation Lien Local handles 720 micro-grant applications per year in three languages. The agent checks workflow entry, inventories and cites documents, applies a published completeness checklist, and prepares a pseudonymized reviewer packet. After approval, it writes and routes the packet. It never scores merit, need, or funding probability.
Consent, known program, channel, readable files
Documents and necessary data only
Administrative facts with page citations
Deterministic completeness checklist
Document request or pseudonymized packet
Sources, transformations, and recipients
Grant system, two reviewers, effect read-back
THE A2 AGENT CARRIES
PEOPLE RETAIN
LOW / CENTRAL / HIGH
60 applications × workflow share × 96 baseline minutes × reduction. Setup is CHF 12,000; recurring cost is CHF 750 per month.
96 → 39 minmedian human time per accepted reviewer packet
×2.46theoretical packet throughput per human hour
58/86reviewer-ready without correction
≈ −39%portfolio ceiling after all 120 applications
100%funding decisions made by people
MISSION BEFORE EFFICIENCY
Telephone, paper, and assisted applications remain available.
The system does not rank vulnerability or infer deservingness.
Rework and stops are reviewed by language, channel, and organization type.
Every decision is explained and can be challenged outside the agent.
Candid: 1% of 529 responding foundations report using GenAI to screen or help decide; 97% say no.
↗FUNCTIONAL ANALOGUE1,000+Degrees of Change handles more than 1,000 applications with 150 volunteer assessors; the provider case describes extraction and staff-reviewed matching, not a causal time result.
↗PROVIDER HIGH BOUND−80%Microsoft reports an 80% reduction in aid-disbursement wait time at NZF. Several changes and a wider automation boundary prevent direct transfer.
↗CASE AUTONOMY DECISION
In this synthetic scenario, 37.5 administrative hours per month are released, about CHF 1,575 net of recurring cost, with a simple payback near 7.6 months. That creates capacity. It does not mean another grant. Any expansion requires affected-person consultation, larger language and channel samples, and a tested challenge path.
WORKED EXAMPLE 04 · PUBLIC SERVICE · PLANNING DOSSIERS · A2
This fictional Swiss municipality tests an A2 agent on administrative completeness, never on planning judgment. The workflow can move faster only if mandate, procurement, evidence, public notice, appeal, archives, and a no-AI service path move with it.
The City of Mont-Rive receives 1,080 planning applications per year in French and German. The agent inventories documents, extracts source-cited administrative facts, and applies a published checklist. After officer approval, it sends, records, and routes. It never declares completeness, interprets law, or recommends approval, refusal, conditions, or priority.
Authority, signature, channel, known request type
Documents, versions, and necessary data
Parcel and project facts with page or plan citations
Published checklist and controlled official sources
Missing-item list or neutral case packet
Sources, uncertainty, recipient, and effects
Send, register, route, and effect read-back
THE A2 AGENT CARRIES
PEOPLE AND THE AUTHORITY RETAIN
LOW / CENTRAL / HIGH
90 applications × workflow share × 145 baseline minutes × reduction. Setup is CHF 48,000; recurring cost is CHF 3,200 per month, including governance and exit.
145 → 58 minmedian human time per accepted case packet
×2.50theoretical packets per officer-hour
121/166approval-ready without correction
≈ −35%portfolio ceiling across all 270 applications
100%public decisions made by qualified people
P0 → P5
Mandate, baseline, non-AI options, and decision authority.
Applicable law, rights, languages, accessibility, and appeal.
Audit, subprocessors, retention, changes, export, and exit.
Representative cases, segments, security, abuse, and outages.
Shadow mode, named approvers, complaint path, and immediate stop.
Signed decision, public notice, archives, fallback, and withdrawal date.
Milton Keynes officers reported faster validation; receipt-to-validation fell from 15.8 to 7.6 days. The three-month supplier case is not causal proof for a whole service.
↗TASK UPPER BOUND18.5h → 16mCambridge reports 16 minutes for PlanAI summaries versus 18.5 human hours on summaries. Planners still read every submission; this narrow ratio cannot price a dossier workflow.
↗GOVERNANCE ANALOGUE6,000+Leeds handles over 6,000 applications yearly and paired six months of co-design with impact reviews, source links, officer approval, an audit trail, and a public ATRS record.
↗P5 · FORMAL PRODUCTION DECISION
The normalized observation is 76.4 hours of monthly capacity, about CHF 2,757 net of recurring cost, and simple payback near 17.4 months. The low case fails the economic gate. Compliance analysis, reasons, prioritization, or a supplier model change returns the system to P0.
WORKED EXAMPLE 05 · COPILOT · INDEPENDENT PROFESSIONAL · A1
A small pilot should answer a small decision. Follow an independent consultant from meeting notes to a reviewed follow-up, without connecting email, calendar, or client systems.
CAMILLE REY · CLIENT FOLLOW-UP
Camille Rey spends a median 44 minutes turning meeting notes into a summary, action list, and follow-up email. The pilot tests a structured first draft while prices, commitments, recipients, and sending remain exclusively human.
−23%median preparation time
12/14ready within 24 hours
4/14major rework
0invented commitments
Confirm 22 historical follow-ups, the manual fallback, data rules, and an eight-hour setup cap.
Tune on 12 authorized cases, then decide on 12 separate cases against thresholds written in advance.
Generate five drafts but reveal them only after the real follow-up has been written manually.
Review nine live drafts against notes. Add commercial content, choose the recipient, and send manually.
CASE BOUNDARY DECISION
The median falls from 44 to 34 minutes and all critical gates pass, but 29% of drafts still need major rework. No automatic sending, full proposal generation, or system connection is justified.
WORKED EXAMPLE 06 · BUSINESS AGENT · INDEPENDENT · A2
Phase 2 keeps the same professional, baseline, and outcome. What changes is the system: approved tools, persistent case state, quality control, execution after approval, and explicit exception handling.
After the A1 copilot pilot, Camille tests a business agent on 20 eligible follow-ups. It reads authorized CRM context, prepares the summary and actions, and checks facts and policy. It waits for one explicit approval before updating the CRM, creating tasks, and sending the reviewed email.
Structured notes and eligibility check
Read-only CRM and client rules
Summary, actions, and follow-up
Facts, dates, policy, and conflicts
One informed human decision
Email, CRM, tasks, and audit log
THE AGENT OWNS
THE PERSON OWNS
44 → 14 minmedian human active time · −68%
×3.1accepted follow-ups per owner-hour
13/20ready to approve without correction
3/20correctly escalated
0unapproved external actions
Drafts one step; the person carries and completes the workflow.
Runs the full bounded workflow and executes only after approval.
Autonomous low-risk sending requires 50 more cases and a new gate.
Use separate identity, least privilege, read-only CRM first, idempotent writes, a kill switch, and a tested manual fallback.
Run 40 representative cases, including conflicts, missing context, price requests, prompt injection, duplicate actions, and unavailable tools.
Compare the complete proposed workflow with the real manual follow-up; no email or write reaches a live system.
Camille reviews one evidence packet, approves or refuses, and the agent executes the authorized actions while logging every effect.
CASE AUTONOMY DECISION
The gain is large because the system now carries the workflow, not because the model merely writes faster. Autonomous sending remains blocked until 50 additional eligible cases show zero critical errors, stable exceptions, no more than 10% major correction, and a verified rollback.
WORKED EXAMPLE 07 · ORCHESTRATED AGENCY · INDEPENDENT · A3
This is where orchestration becomes useful: the work contains distinct research, analysis, quality, and execution roles that can run in parallel. The scope remains one eligible service, not the whole business.
Camille delivers a standardized operational diagnostic for existing small-business clients. After the client interview, the agency qualifies the case, retrieves authorized evidence, scores the process, produces the report and action plan, challenges its own conclusions, then performs low-risk CRM, task, delivery, and scheduling actions inside a pre-approved service policy.
Assigns work, enforces the case policy, resolves dependencies, stops on disagreement, and accepts no specialist’s self-reported success without effect evidence.
Identity, eligibility, minimization
Authorized sources and traceable citations
Diagnosis, scoring, and uncertainty
Report, actions, and client-ready structure
Facts, contradictions, risk, and permissions
Delivery, CRM, tasks, and scheduling
LIKE-FOR-LIKE BENCHMARK
×7.9accepted diagnostics per owner-hour
9/12accepted without major rework
8/12eligible cases completed straight through
4/12stopped and escalated before effect
5h 20median internal cycle vs 18h
0unauthorized commitments or writes
Separate roles, inputs, outputs, permissions, failure boundaries, effect evidence, and the situations that must remain human.
Run 60 cases through manual, copilot, single-agent, and orchestrated conditions; measure accepted output, not agent activity.
Run specialists in parallel without live effects. Inject disagreement, stale memory, tool outages, duplicated events, and poisoned source content.
Allow only catalogued low-risk effects on eligible cases. Guardian veto, cost limits, rollback, and human escalation remain active.
CASE SCOPE DECISION
The agency may continue for the defined diagnostic and its catalogued effects. It may not choose new services, prices, contracts, clients, data classes, or permissions. Broad multi-system autonomy requires a separate mandate, an independent audit, and evidence across several workflows.
TECHNICAL PROGRESSION
Move one level at a time. Stop when a simpler system meets the need.
QUICK CONTROL ORIENTATION
This is internal triage, not a legal classification.
VERSIONED CONTROL CROSSWALK · JSON 1.1
Organization, impact, autonomy, use pattern, and jurisdiction resolve to stable control IDs. Each row exposes its trigger, evidence, decision gates, lifecycle phases, and dated source references.
20candidate controls
Independentorganization
R1 × A2R1 · assistance with review · A2 · action after explicit approval
Retrievaluse pattern
Switzerland + EUjurisdiction route
12versioned sources
Tie the system to one measurable problem, a bounded mandate, prohibited actions, and a person accountable for the final decision.
Make official and informal AI use visible, with ownership, provider, model, risk, autonomy, and review dates.
Compare the AI system with the real current process using complete denominators and a business or public-service outcome.
Prevent a low technical autonomy label from hiding high human impact, sensitive data, or regulated responsibility.
Record provenance, purpose, access, processing locations, reuse, retention, deletion, and subprocessors before using personal or confidential data.
Trigger conditionWhen personal or confidential data is processed
Preserve audit, change notice, export, deletion, continuity, and exit rights when an external provider is part of the system.
Trigger conditionWhen an external provider is used
Fix the decision unit, test population, thresholds, critical errors, and authorized judge before optimizing or piloting.
Keep a decision set separate from development and reject averages that conceal failure on a critical segment or an unacceptable error.
Trigger conditionWhen people may be materially affected
Use attributable identities, read-only defaults, temporary permissions, approved destinations, isolated secrets, and enforceable limits.
Trigger conditionWhen external content or tools cross a trust boundary
Test hostile content, identity confusion, poisoned context, unsafe tool parameters, partial failure, loops, data leakage, and attempts to disable controls.
Trigger conditionWhen external content or tools cross a trust boundary
For every agent action, bind the tool, permission, approval, idempotency key, destination, effect read-back, rollback, and owner.
Trigger conditionWhen the system can create an external or irreversible effect
Ensure a qualified person has enough information, time, authority, and an operational channel to reject, correct, override, and hear appeals.
Trigger conditionWhen people may be materially affected
Link each case to the exact configuration, evidence, approval, tool parameters, external effect, read-back, actor, and time.
Trigger conditionWhen people may be materially affected
Authorize only the proven scope and maintain tested suspension, safe routing, evidence preservation, reconciliation, rollback, and re-authorization steps.
Trigger conditionWhen the system can create an external or irreversible effect
Give a reachable owner the authority and tested procedure to contain, preserve evidence, notify, revoke access, recover, and trigger reassessment.
Trigger conditionWhen external content or tools cross a trust boundary
Monitor business outcomes, critical segments, corrections, incidents, drift, and cost; return to the right gate after any material model, data, tool, permission, or population change.
Stop the system cleanly by revoking access, exporting what continuity requires, disposing of data, preserving required evidence, and restoring a safe process.
Record task, interaction, knowledge, modality, deployment, operating mode, effect, and jurisdiction separately from impact and autonomy before selecting a system.
Add retrieval, classification, prediction, conversation, multimodal, or agentic measures and stop rules whenever those patterns are present.
Extend the threat model for retrieval, predictive models, conversation, multimodal inputs, code execution, persistent memory, and inter-agent communication as applicable.
USE THE PLAYBOOK
Copy the operational templates, complete the first gate, and keep the evidence with the project.
Owner, baseline, outcome, and boundaries.
↗02Value and difficulty kept separate.
↗03Scenarios, controls, and residual risk.
↗04Metrics, segments, thresholds, and stop rules.
↗05Value, reliability, and risk judged separately.
↗06Contain, qualify, recover, and learn.
↗07Complete tasks, assistive technologies, and an equivalent non-AI channel.
↗08Affected groups, rights, safeguards, recourse, and residual impact.
↗09Observed evidence, anonymization review, and explicit transfer limits.
↗