REALISTIC VALUE
What ROI can a business AI agent realistically deliver?
There is no defensible universal low and high return for a business AI agent. The strongest direct field experiment found 16.8% faster eligible chats, but only 5.8% of chats were eligible, so the effect across all chats was 3.2% and customer ratings fell on eligible chats. Enter a low and high hypothesis for your workflow, apply it to the baseline human hours carried by eligible work, then count supervision, exceptions, incidents, maintenance, and model costs.
Reading key: A0 to A4 describe the system’s authority. R0 to R3 describe possible impact. Letters A to E, when attached to a source, describe evidence strength only.
Key takeaways
Key takeaways
- 01
Apply the gain only to eligible work, never to the whole workload by default.
- 02
Count corrections, approvals, exceptions, and maintenance as real labor.
- 03
Separate released capacity from cash savings and new revenue.
Three different results, not one universal range
| Evidence | Measured setting | Observed result | Correct interpretation |
|---|---|---|---|
| Small-business RCT | Open-ended business advice | 0% average; about −8% to +15% by baseline skill | Judgment and implementation determine the sign |
| Business-agent RCT | Standardized customer-service chats | −16.8% eligible duration; −3.2% across all chats | Eligibility and quality limit the headline gain |
| Agency benchmark | 46 concurrent simulated corporate tasks | 15.2% completed; 3.5x relative to 4.3% | Capability signal, not a production ROI |
Count the complete workload
If eligible cases account for 70% of baseline human hours and the agent reduces that time by 60%, the gross reduction across the complete workload is 42% before operating costs. A share of cases is not enough: measure the hours in eligible cases and in the complete workload. A denominator is simply the total you compare against. Keep the same workload before and after AI. Hours saved are not automatically money saved; you still decide how to use that time.
This prevents a common reporting error: publishing a strong rate on accepted cases while hiding the cases that never entered the system.
Measure an outcome that the organization values
Time saved is useful only when it changes capacity, cost, service, quality, or revenue. Record accepted cases per owner-hour, major corrections, critical errors, cost per accepted result, and the downstream result. Keep each measure separate.
WORKED EXAMPLE
Example: 40 monthly cases
At 60 manual minutes per case and 70% assumed eligibility, editable challenge hypotheses of 20% and 50% produce 5.6 to 14 gross hours per month. A 40-hour setup would take about 2.9 to 7.1 months to absorb before recurring costs. These are scenario inputs, not an evidence range; the pilot must replace them with observations.
- 28 assumed eligible cases
- 5.6 to 14 gross hours
- 2.9 to 7.1 months before recurring costs
Sources and limits
Sources and limits
These sources bound the answer. They do not turn one published case into a promise for your organization.
- 01Agentic AI in customer service ↗
Direct randomized field evidence for a bounded business agent.
- 02The Uneven Impact of Generative AI ↗
Randomized small-business evidence with an average null and heterogeneous effects.
- 03Writing Code vs. Shipping Code ↗
Activity gains attenuate at projects, releases, and downstream usage.