METR's February 2026 update lists the reasons its productivity experiment stopped working. I have run into every one of them inside a real team.
The study randomized tasks into AI allowed and AI disallowed conditions for 57 experienced open source developers, across 143 repositories. The estimate for returning developers moved from 19% slower in 2025 to 18% faster, with wide intervals. METR trusts the direction more than the size: developers withheld 30% to 50% of tasks rather than do them by hand, they chose different kinds of tasks when an agent was available, output quality differed between conditions, and time became hard to log because people worked on something else while the agent ran.
Those four effects are exactly what changes in a team once adoption is real. Task selection shifts toward what the agent does well. Definition of done shifts because the agent writes the tests you used to skip. The developer's clock stops being the unit of work.
Gergely Orosz described the other half on X: engineering teams "patting themselves on the back" on velocity while power users notice small regressions nobody reports.
Which is why I no longer ask whether a developer is faster with AI. I ask what the team shipped, what came back, and how long review took.
Sources
- METR, We are Changing our Developer Productivity Experiment Design, February 2026. https://metr.org/blog/2026-02-24-uplift-update/
- Gergely Orosz on X. https://x.com/GergelyOrosz/status/2098329680461885884
