FROM DEMO TO EVIDENCE
How do you pilot a business AI agent?
For this pilot with human approval, work in four stages: write the test rules and measure current work, test the same set of requests, run the system without changing real tools, then allow the planned actions on a few cases after explicit approval. Set success and failure rules for value, quality, safety, action records, and eligible cases before starting. A critical unauthorized effect must stop the pilot regardless of average productivity.
Reading key: A0 to A4 describe the system’s authority. R0 to R3 describe possible impact. Letters A to E, when attached to a source, describe evidence strength only.
Key takeaways
Key takeaways
- 01
Write the pass, rework, and stop decisions before observing results.
- 02
Test failures, duplicates, abstentions, and rollback, not only happy paths.
- 03
Increase autonomy only after the current level passes on real evidence.
Four pilot stages
| Stage | System effect | Evidence | Exit condition |
|---|---|---|---|
| Baseline | None | Manual time, quality, volume, eligibility | Decision and thresholds frozen |
| Offline | None | Fixed real cases and adversarial cases | No critical failure |
| Shadow | No external effect | Full workflow traces | Stable accepted quality |
| Bounded live | Approved reversible effects | Outcomes, exceptions, rollback | Pass, rework, or stop |
Define five separate gates
Value asks whether accepted work improves. Quality asks whether outputs meet the review standard. Safety asks whether any critical or unauthorized effect occurred. Traceability asks whether approvals and effects can be reconstructed. Eligibility shows how much of the real workload the result covers. The case counts and duration below describe this example, not a universal sample size. Before testing, name what a successful reply looks like and who stops the system after an unsafe action.
Treat missing evidence as unknown
An incomplete trace or a sample that never reached the required size is not a pass and not automatically a failure. It means the pilot cannot support the decision yet. Extend the observation or repair the evidence path without increasing autonomy.
WORKED EXAMPLE
A practical A2 floor
For a bounded business agent, start with about 40 frozen cases and 20 eligible live cases over at least 30 days. Increase the sample for rare failures, protected groups, high variance, or higher-stakes effects. Calendar time never replaces enough eligible cases.
- 40 frozen cases
- 20 bounded live cases
- Zero critical unauthorized effect
Sources and limits
Sources and limits
These sources bound the answer. They do not turn one published case into a promise for your organization.
- 01NIST AI Risk Management Framework ↗
Roles, measurement, risk treatment, and lifecycle governance.
- 02NIST Generative AI Profile ↗
Generative AI risk actions and evidence considerations.
- 03OWASP GenAI Security Project ↗
Current security guidance for LLM and agentic systems.