Skip to content
← Back to blog
Research·September 17, 2026·7 min read

METR's randomized trial found AI made developers 19% slower. Its new survey found people claiming a 3x speed gain.

METR's May 2026 survey found workers self-report up to a 3x AI speed gain. Its 2025 randomized trial found developers 19% slower. Same lab, opposite numbers.

METR ran the randomized controlled trial that found AI coding tools made experienced developers 19% slower, not faster, while those same developers believed they'd sped up by 20%. A year later, the same organization asked people to self-report the number directly, no stopwatch, no control group, and got a median of 1.4 to 2x more value from their work and a 3x jump in speed. METR published both studies. The lab that proved self-reported speed gains can run backwards followed it with a survey built entirely out of self-reported speed gains, and then told its own readers not to trust it either.

What the new survey actually found

METR's Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity, published May 11, 2026, fielded responses from 349 technical workers between February and April 2026: 87 software engineers, 71 researchers, 129 academics and PhD students, and 48 founders and managers, recruited through GitHub, institutional directories, METR's own staff networks and X, with roughly a 2% response rate from the email outreach arm. About 70% of respondents were paid, averaging around $200 each.

The headline numbers: a median self-reported 1.4 to 2x change in "the value of their work" due to AI tools, and a median 3x change in speed. Asked to look backward, the same group put their March 2025 value multiplier at 1.3x. Asked to look forward, their March 2027 forecast climbed to 2.5x. Read as a single line, it's a tidy hockey stick: AI's contribution to a technical worker's output roughly doubling every year, self-reported by the people doing the work.

METR doesn't trust its own numbers, and says so

Here's the part that should stop a reader before they act on that hockey stick: METR wrote the caveat into the same post. "Survey results are not necessarily grounded in reality," the write-up states, pointing to its own earlier finding that people overestimated AI's effect on their task time by 40 percentage points on average. When METR's researchers dug into seven of the most extreme responses, ones claiming a 10x or greater change in value, their read was that "participants are overstating their change in value." A research group ran the survey, published the survey, and then published its own doubt about the survey, in the same document. That's unusual candor, and it's also the whole story: a number this easy to produce (an opinion, typed into a form) is exactly the number a team under pressure to justify a tooling budget will reach for first.

The same lab already ran the controlled version

METR didn't arrive at that skepticism from nowhere. Its July 2025 paper, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, randomly assigned AI or no-AI to 246 real tasks across 16 experienced developers working in codebases they knew intimately, the same design a drug trial uses to separate a real effect from a placebo. The result: allowing AI increased completion time by 19%. The developers doing the tasks believed AI had made them 20% faster, a 39-point gap between what they felt and what a clock measured. We covered that study in detail here, and it's the reason a survey full of self-reported multipliers, from the same organization, deserves a second look rather than a headline.

Two studies, two very different instruments

July 2025 randomized trialMay 2026 self-report survey
MethodAI/no-AI randomly assigned per taskParticipants estimate their own change in value and speed
Sample16 experienced developers, 246 real tasks349 technical workers, 87 of them software engineers
Ground truthMeasured task completion timeNone; entirely self-reported
Headline result19% slower with AI1.4-2x more value, 3x faster, self-reported
METR's own caveatNone needed; time was measured directly"Not necessarily grounded in reality"

The gap between these two rows isn't a contradiction so much as a demonstration of what each method can and can't tell you. A stopwatch can't be talked into feeling generous. A form field can.

A worked example

Picture a 70-person, Series B developer-tools company six weeks into a company-wide pilot of an agentic coding assistant. The VP Eng sends out a short internal survey asking engineers to estimate how much faster the tool has made them. The median answer comes back around 2x, a few enthusiastic outliers claim 5x, and the VP Eng takes that number into the next budget cycle to justify a seat-based contract across the entire 45-person engineering org, no pilot cohort held back for comparison. Nobody pulled a single cycle-time number from the ticketing system first. That's the exact shape of evidence METR's own paper warns against trusting, run at a company with no research team dedicated to catching it.

What this means for a VP Eng deciding on tooling or headcount

A self-reported number is worth exactly one thing: a read on sentiment and adoption, which matters for change management and nothing else. It is not evidence for a budget decision, and treating it as one is the same mistake whether it comes from an internal survey or a vendor's case study built the same way. If your team's confidence in an AI tool has to become a number leadership acts on, run the cheap version of METR's design instead: hold out a comparison group, or compare similar tickets before and after adoption using cycle time and review-cycle count already sitting in your own systems, not a form asking someone to guess. That's a genuine opinion worth defending on a sales call, not a hedge: a survey result should change how you talk to your team, and it should not, by itself, move a dollar of engineering spend.

The honest exception cuts the other way. If the decision on the table is whether developers feel the tool is worth keeping, not whether it's actually shortening delivery, a self-reported survey is the right instrument for that narrower question, and building a randomized trial to measure morale would be overkill. The mistake isn't asking the question. It's answering a productivity question with a morale-shaped instrument.

A founder inside a 30-person company usually can't run a 16-developer randomized trial before the next roadmap review is due, and METR's own paper doesn't pretend the option is universal. What transfers at any size is the discipline behind it: pick one metric a stopwatch or a ticketing system already produces (cycle time, review-to-merge duration, defect rate on a release), compare it across a real before-and-after period or a held-out group, and treat anyone's felt-speed estimate, including your own, as a hypothesis to check against that number rather than a finding to act on.

If your team is weighing a bigger AI-tooling commitment, or a headcount plan, off internal sentiment rather than a number you've actually measured, that's worth a second look before the contract renews on the strength of a survey. Our engineering team builds the lightweight evaluation harnesses that turn "it feels faster" into a number worth trusting, the same discipline our guide to building an eval harness for AI features argues has to exist before any AI investment scales past a pilot. Our Silicon Valley team can help you design that comparison before the next budget cycle, and telling us what you're measuring costs less than a year of seats bought on a number nobody checked.

Sources

Frequently asked questions.

METR surveyed 349 technical workers between February and April 2026 and found a median self-reported change of 1.4 to 2x in the value of their work and 3x in speed due to AI tools. The same group retrospectively estimated a 1.3x value change for March 2025 and forecast 2.5x for March 2027.

It sits in tension with it rather than replacing it. METR's July 2025 randomized controlled trial measured actual completion time across 246 tasks and found AI made 16 experienced developers 19% slower, even though those same developers believed they were 20% faster. The 2026 survey asked for a self-reported estimate with no measured baseline, which is a different kind of evidence.

METR's own write-up states that "survey results are not necessarily grounded in reality" and cites its earlier finding that people overestimated AI's effect on their task time by 40 percentage points on average. Reviewing the most extreme responses in the 2026 survey, claims of a 10x or greater change, METR's own researchers concluded participants were likely overstating the change.

Compare a metric already produced by existing systems, such as cycle time, review-to-merge duration or defect rate, across a real before-and-after period or a held-out comparison group, rather than asking engineers to estimate their own speed change. That mirrors the design METR itself used in its measured randomized trial rather than its self-report survey.

Yes, for a narrower question than ROI: gauging team sentiment or adoption of a new tool, where the goal is understanding how people feel about it rather than proving it saved time. It becomes the wrong tool the moment its number is used to justify a budget or headcount decision without a measured baseline behind it.