# What recent implementations can offer you

Targeted review on 5 September 2026. These cases complement timed studies; they do not replace them.

You can reuse a working method and try your own values in the calibrator. An internal estimate, elapsed time and a quality score are not interchangeable measures of human time saved.

## Document checks across 41 files

Legora · 2026-09-03 · [Original source](https://openai.com/index/legora-financial-statement-review-with-astra/)

Legora reports checking 41 documents in minutes, with experts making the final decisions. Its 40% improvement is a benchmark score, not time saved; it does not enter the time calculation.

**What you can transfer :** Compare figures against supporting documents and record discrepancies for review.

**Limits of the figure :** No matched manual baseline; 40% concerns one workflow versus about 3% across the benchmark.

## Fewer manual fixes in visual prototypes

Playco · 2026-09-03 · [Original source](https://openai.com/index/playco-game-prototyping-with-astra/)

Playco reports half as many manual fixes when building game prototypes with a newer model. This is not half the total work time and does not enter the time calculation.

**What you can transfer :** Generate a prototype, inspect it in its real tool, test it, then correct it.

**Limits of the figure :** Three prototypes; supplier-published case; counts of fixes do not measure their duration.

## Marketing launch with connected agents

Stampli · 2026-08-20 · [Original source](https://openai.com/index/stampli/)

Stampli estimates 243 hours without AI versus about 77 with AI for launch production, with human approval. The manual baseline was modelled, so this does not enter the calculation automatically.

**What you can transfer :** Connect product context, draft launch assets, then review and approve each deliverable.

**Limits of the figure :** Estimated baseline, not two equivalent launches timed with and without AI. Use as a named planning example.

## Technical tasks with a copilot or an agent

Anthropic · 2026-06-18 · [Original source](https://www.anthropic.com/research/project-fetch-phase-two)

On four shared tasks, Anthropic reports 361 minutes without Claude, 181 with assistance and about 9.6 with the newer agent. These elapsed times do not enter the human-time calculation.

**What you can transfer :** Connect documented tools, write integration code and verify the result under human authorization.

**Limits of the figure :** Historical human comparison; three agent trials; some physical tasks excluded. Researcher approves commands.

## Coding-agent adoption in real teams

Microsoft · 2026-07-01 · [Original source](https://arxiv.org/html/2607.01418v1)

Microsoft estimates about 24% more merged code changes with CLI agents over four months. This measures delivery volume, not saved minutes, and does not enter the time calculation.

**What you can transfer :** Use connected coding tools and peer examples to improve delivery on existing repositories.

**Limits of the figure :** Observational comparison with synthetic controls; other AI tools already available. A merged change is not business value.

## Generate, test and select better solutions

Google DeepMind · 2026-05-07 · [Original source](https://deepmind.google/blog/alphaevolve-impact/)

DeepMind reports better algorithms, including fewer detection errors and more feasible grid solutions. These quality gains do not enter the time calculation; the useful pattern is generate, test and select.

**What you can transfer :** Generate candidate algorithms and select them using an explicit automated evaluator.

**Limits of the figure :** Scientific and infrastructure outcomes, not measured human-time savings. Local tests must define what a good solution means.

## Parallel research on a measurable objective

Anthropic Fellows · 2026-08 · [Original source](https://alignment.anthropic.com/2026/automated-alignment-researchers/)

Automated researchers improved ten tested failure categories; their best methods beat human ideas in this protocol. The result is not a human-time saving and does not enter the calculation.

**What you can transfer :** Propose alternatives in parallel, evaluate separately and keep validated improvements.

**Limits of the figure :** Selected best methods, substantial compute and review, bounded objectives. No universal research productivity ratio.

## Engineering done, research question unresolved

Independent research authors · 2026-08-07 · [Original source](https://arxiv.org/abs/2607.27191)

In two research cases, agents completed the engineering but did not answer the central research questions well enough. This capability test does not enter the time calculation.

**What you can transfer :** Separate technical execution from the expert judgment that makes the result useful.

**Limits of the figure :** Two cases, six days and substantial compute; neither a general failure rate nor a time-saving estimate.

## Parallel agents beyond software teams

OpenAI · 2026-06-25 · [Original source](https://openai.com/index/how-agents-are-transforming-work/)

OpenAI describes delegated work across technical and non-technical teams. Parallel agent hours are not human hours saved, so these usage figures do not enter the time calculation.

**What you can transfer :** Delegate separate tasks with shared context, explicit outputs and human review.

**Limits of the figure :** Usage telemetry and estimated task durations, not a matched measure of accepted human work.

## Voice assistant connected to business tools

xAI · 2026-07-01 · [Original source](https://x.ai/news/grok-voice-agent-builder)

xAI combines calls, document lookup and business tools in a voice-agent builder. Fast configuration is not a validated rollout or measured saving, so it does not enter the time calculation.

**What you can transfer :** Answer a caller from approved documents, perform permitted actions and hand off exceptions.

**Limits of the figure :** Product description and provider benchmark. Test noise, interruptions, accents, consent and actual effects locally.

These ten cases are available in the registry by task. Their context-only status prevents automatic conversion into a savings percentage; you can still simulate the mechanism with clearly labelled assumptions.

[Calculation method](../docs/task-time-evidence.md)
