Back to Blog
Productivity2026-08-098 min read

Developers Thought AI Made Them 20% Faster. It Made Them 19% Slower.

In 2025, METR ran a randomized controlled trial on sixteen experienced open-source developers. These were not students. They averaged five years working on mature repositories, projects with tens of thousands of stars and over a million lines of code. Each contributed real issues from their own backlog: bugs, features, refactors, the actual work.

Every task was randomly assigned to allow AI tools or forbid them. Mostly Cursor Pro with Claude.

Before starting, the developers predicted AI would make them 24% faster.

Afterwards, having done the work, they estimated it had made them 20% faster.

They were 19% slower.

Read that sequence again, because the interesting part is not the slowdown. It is that the estimate barely moved after they had lived through it. Direct personal experience of being slower produced almost no correction. A roughly forty percentage point gap between what these people believed about their own hours and what those hours actually were.

How anyone found out

Here is the part of the study that we find hard to get past.

METR could not establish what happened by asking. They had already asked, and got a number that was wrong by 39 points. So they recorded the developers' screens and manually analysed 128 screen recordings, 143 hours of video, to work out where the time had actually gone.

That is the cost of the ground truth. A hundred and forty-three hours of humans watching video, because self-report was worthless and nothing else was measuring.

Every knowledge worker on earth is running the same experiment on themselves, continuously, with no recordings and no control group.

This is not a story about AI

It is easy to read the METR result as an anti-AI finding, and the follow-up work suggests it is not that.

METR's early-2026 survey of technical workers found a median self-reported speedup of 3x from AI. But when the same people were asked about value delivered rather than time saved, the median was 1.4x to 2x. The researchers are explicit that speed estimates run high relative to value, because people drift toward tasks AI does quickly regardless of whether those tasks are worth doing.

They are equally blunt about their own data: survey results "are not necessarily grounded in reality," and they cite their earlier finding that people misjudged AI's effect on their time by 40 percentage points on average.

So the finding underneath the finding is simpler and much older than AI:

People cannot estimate where their own time goes. Not roughly. Not within 40 points.

AI just made the error easier to see, because it introduced a big, sudden change to working patterns and gave everyone a strong opinion about it.

Why we are so bad at this

A few reasons, all of which compound.

Effort is not duration. The forty minutes of hard concentration feel long. The ninety minutes of waiting, re-prompting, reading a diff, getting distracted and coming back feel like twenty. We remember exertion, not elapsed time.

Interruptions are invisible in memory. Nobody encodes the eleven Slack context switches. The narrative that survives is "I worked on the migration this morning," which is technically true and useless.

Vivid moments dominate. One brilliant AI completion that saved an hour is memorable. Forty minutes spread across three failed prompts is not, so the average gets pulled toward the highlight.

We are motivated. Nobody wants to conclude that the expensive new tool did nothing. Nobody wants to conclude that Tuesday disappeared.

The 2026 version of the problem

There is a newer wrinkle that makes this worse, and it is specific to how people work now.

METR notes that time-on-task measurements are unreliable for the growing fraction of developers running multiple AI agents concurrently. If you kick off three agents and supervise them in rotation, what does "time spent on task A" mean? You were present for all three and working on none of them in the way that word used to imply.

Traditional output metrics have degraded for the same reason. Pull requests per week and lines of code were always crude, and AI-assisted workflows inflate the volume without necessarily moving the value, which is why engineering organizations spent 2026 layering DORA, SPACE and DX Core 4 on top of each other trying to triangulate something real.

Meanwhile the individual question stayed exactly the same and got no easier: where did my week actually go?

What to do about it

You cannot introspect your way to the answer. That is the entire lesson of the METR result, and it is the one part that generalizes to everybody, whether or not you write code.

You have three options.

1. Manual time tracking. Start a timer, stop a timer, tag it. Accurate when you do it, and you will not do it. The discipline required is precisely the discipline the people who need it most do not have, and the act of logging is itself an interruption.

2. Record everything. Screenshot your display continuously and index it, the way METR did with video, which is now a product category. It works, and it means creating a complete searchable record of your entire digital life on your disk. We looked at what that trade actually involves.

3. Measure automatically, at low resolution. Have something watch which app and window you are in, decide what that meant, and keep only the conclusion. No timers to start, no recordings to protect.

The third is what we built Cronus to do.

Cronus Interface

It reads the app you are in, the window title and the browser URL, and an AI model works out what that activity was and whether it served what you told it you were working on. It stores the interval and the meaning, not the contents. No screenshots, ever. Sensitive information is redacted on your device before anything is transmitted, and we are GDPR compliant. Three days free, then $6 a month.

The uncomfortable part

If you install something like this, you should expect the first week to be unpleasant, and you should not treat that as a malfunction.

The METR developers were off by forty points about a thing they had just personally experienced. Your estimate of last Wednesday is not better than theirs. The number you get back will very likely be lower than the number in your head, and the shape of the day will be worse than you remember: more fragmented, more scattered, with the deep work compressed into a much smaller window than it feels like.

That gap is the entire value. You cannot fix a forty-point error you cannot see, and the one thing we know for certain is that thinking harder about it does not close it.

Those developers had the strongest possible evidence available, their own lived experience of doing the work, and it moved their estimate by four points.

Try Cronus. Find out what your week actually was.


Sources: METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" and METR, "Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity".

Find out where your week actually goes

The 5-day time audit: one short email a day, walking you through the experiments in this post. You will end the week with a real answer instead of an impression. Free, no app required, and it ends after five days.