Evidence and references
What recent implementations can offer you
Targeted review on 5 September 2026. These cases complement timed studies; they do not replace them.
You can reuse a working method and try your own values in the calibrator. An internal estimate, elapsed time and a quality score are not interchangeable measures of human time saved.
Document checks across 41 files
Legora · 2026-09-03 · Original source
Legora reports checking 41 documents in minutes, with experts making the final decisions. Its 40% improvement is a benchmark score, not time saved; it does not enter the time calculation.
What you can transfer : Compare figures against supporting documents and record discrepancies for review.
Limits of the figure : No matched manual baseline; 40% concerns one workflow versus about 3% across the benchmark.
Fewer manual fixes in visual prototypes
Playco · 2026-09-03 · Original source
Playco reports half as many manual fixes when building game prototypes with a newer model. This is not half the total work time and does not enter the time calculation.
What you can transfer : Generate a prototype, inspect it in its real tool, test it, then correct it.
Limits of the figure : Three prototypes; supplier-published case; counts of fixes do not measure their duration.
Marketing launch with connected agents
Stampli · 2026-08-20 · Original source
Stampli estimates 243 hours without AI versus about 77 with AI for launch production, with human approval. The manual baseline was modelled, so this does not enter the calculation automatically.
What you can transfer : Connect product context, draft launch assets, then review and approve each deliverable.
Limits of the figure : Estimated baseline, not two equivalent launches timed with and without AI. Use as a named planning example.
Technical tasks with a copilot or an agent
Anthropic · 2026-06-18 · Original source
On four shared tasks, Anthropic reports 361 minutes without Claude, 181 with assistance and about 9.6 with the newer agent. These elapsed times do not enter the human-time calculation.
What you can transfer : Connect documented tools, write integration code and verify the result under human authorization.
Limits of the figure : Historical human comparison; three agent trials; some physical tasks excluded. Researcher approves commands.
Coding-agent adoption in real teams
Microsoft · 2026-07-01 · Original source
Microsoft estimates about 24% more merged code changes with CLI agents over four months. This measures delivery volume, not saved minutes, and does not enter the time calculation.
What you can transfer : Use connected coding tools and peer examples to improve delivery on existing repositories.
Limits of the figure : Observational comparison with synthetic controls; other AI tools already available. A merged change is not business value.
Generate, test and select better solutions
Google DeepMind · 2026-05-07 · Original source
DeepMind reports better algorithms, including fewer detection errors and more feasible grid solutions. These quality gains do not enter the time calculation; the useful pattern is generate, test and select.
What you can transfer : Generate candidate algorithms and select them using an explicit automated evaluator.
Limits of the figure : Scientific and infrastructure outcomes, not measured human-time savings. Local tests must define what a good solution means.
Parallel research on a measurable objective
Anthropic Fellows · 2026-08 · Original source
Automated researchers improved ten tested failure categories; their best methods beat human ideas in this protocol. The result is not a human-time saving and does not enter the calculation.
What you can transfer : Propose alternatives in parallel, evaluate separately and keep validated improvements.
Limits of the figure : Selected best methods, substantial compute and review, bounded objectives. No universal research productivity ratio.
Engineering done, research question unresolved
Independent research authors · 2026-08-07 · Original source
In two research cases, agents completed the engineering but did not answer the central research questions well enough. This capability test does not enter the time calculation.
What you can transfer : Separate technical execution from the expert judgment that makes the result useful.
Limits of the figure : Two cases, six days and substantial compute; neither a general failure rate nor a time-saving estimate.
Parallel agents beyond software teams
OpenAI · 2026-06-25 · Original source
OpenAI describes delegated work across technical and non-technical teams. Parallel agent hours are not human hours saved, so these usage figures do not enter the time calculation.
What you can transfer : Delegate separate tasks with shared context, explicit outputs and human review.
Limits of the figure : Usage telemetry and estimated task durations, not a matched measure of accepted human work.
Voice assistant connected to business tools
xAI · 2026-07-01 · Original source
xAI combines calls, document lookup and business tools in a voice-agent builder. Fast configuration is not a validated rollout or measured saving, so it does not enter the time calculation.
What you can transfer : Answer a caller from approved documents, perform permitted actions and hand off exceptions.
Limits of the figure : Product description and provider benchmark. Test noise, interruptions, accents, consent and actual effects locally.
These ten cases are available in the registry by task. Their context-only status prevents automatic conversion into a savings percentage; you can still simulate the mechanism with clearly labelled assumptions.
To print or save as PDF: Ctrl+P (⌘P on Mac).
Source and history · GitHub