AI & TECHNOLOGY / RESEARCH INTO PRACTICE
AI Productivity: Start the Clock Before the Prompt
An answer appears in 30 seconds. Twenty minutes later, you are still fixing it. To measure AI productivity, the finish line has to be usable work, not the first fluent response.
Your practical takeaway: Compare similar tasks to the same quality standard, counting prompting, checking and repairs. Track saved time separately from the value of extra work.
In this guide
A large gain, with an important label
In a report published May 11, 2026, METR surveyed 349 technical workers. Across three value questions, median self-reported multipliers were 1.4× to 2×. The median reported speed multiplier was 3×, meaning participants estimated that comparable work took about one-third as long. These were self-reports, not a stopwatch experiment demonstrating those gains. The sample was recruited by convenience, and METR warns about selection and inaccurate counterfactual estimates. Read the survey and methodology.
That leaves two different questions on your desk: did you finish comparable work sooner, and did you produce something more valuable? A dashboard nobody uses can raise a speed score and leave the project’s value unchanged.
Why the old “AI made developers slower” headline needs a date
METR’s earlier experiment found that early-2025 AI tools increased completion time by 19% for experienced developers working in familiar repositories. In its February 24, 2026 update, METR said newer data gave an unreliable estimate: some developers and tasks were being selected out because participants did not want to work without AI, and parallel agents complicated time measurement. The researchers thought newer tools likely helped more, while stressing that the evidence for the size of that change was weak. Read the experimental update.
Neither headline answers how much time your current workflow saves. Use the papers to notice measurement problems, then design a small check around work you actually do.
One answer, four parts of the bill
Hypothetical example: a short research brief takes 24 minutes without AI and meets a written acceptance checklist. A similar brief with AI has the following log. These figures demonstrate the calculation; they are not results from METR or our own experiment.
| Activity | Minutes with AI |
|---|---|
| Prepare the prompt and inputs | 3 |
| Wait for a draft | 0.5 |
| Open sources and check claims | 8 |
| Repair omissions and format the accepted brief | 7 |
| Total | 18.5 |
The saving is 24 − 18.5 = 5.5 minutes. As a share of the original time, 5.5 ÷ 24 × 100 is about 22.9% less time. The task speed ratio is 24 ÷ 18.5, or about 1.30 times. Those are two ways to describe the same example, not two separate benefits. If repairs instead take 15 minutes, total time becomes 26.5 minutes and the apparent shortcut becomes a 2.5-minute overrun.
Make the comparison hard to fool
Write the acceptance checklist before either attempt. For a brief, it might require every material claim to have an opened source, the requested scope to be covered, and the final file to be ready to use. Keep that standard identical. A shorter answer with missing evidence has not won.
Compare several tasks of similar difficulty. Alternate which workflow you use first, because fatigue and practice can influence later attempts. Avoid doing the identical task twice and counting the second run as a fair test: you already learned its answer. Record topic, difficulty, tool and model version, active work time, elapsed time and repair time.
If an agent runs while you do another job, keep active work and elapsed time in separate columns. Do not count the waiting period as fully occupied labour, or claim the same saved minutes twice. Record failed attempts as well as successes, including the ones you abandon and redo manually.
What if AI changes the work you choose?
A timer cannot tell you whether an extra chart, prototype or ten-page memo was worth producing. Add one sentence to the log: “This output helped ___ make ___ decision.” If you cannot fill it in, mark its value as unproven rather than converting its volume into a productivity gain.
That is our proposed workflow check. It is more modest than a causal field experiment, but it can reveal where your own repair bill is hiding. The AI model comparison guide provides a broader scoring rubric when you want to compare tools.
Choose the next task from the log
For tasks with repeatable requirements and short repair times, try AI again. For tasks where checking dominates, investigate why: missing source access, vague inputs, or a job that requires specialist judgement. Change one part of the workflow and measure another comparable task. Do not assume that a bigger prompt fixes the problem.
Keep the log for one working week. At the end, total only accepted outputs. Your decision then rests on the time and usefulness you recorded, with the failed runs still visible.
Primary sources and research dates
- METR, May 11, 2026: self-reported productivity survey
- METR, February 24, 2026: experimental update and selection effects
Research summaries above are linked to their original sources. Worked examples and practice methods are our own and are labelled separately from study results. This article was written with AI assistance; source claims and calculations were checked before publication.