Every few weeks a screenshot goes around claiming an AI agent just worked unsupervised for a full day. The organisation that actually measures this published its own updated numbers on January 29, 2026, and they say something more modest: the strongest model in that report can reliably, meaning a 50% success rate, finish a task that takes a skilled human about five hours, not eight and nowhere near sixteen. The gap between the screenshot and the study is the whole story, because the number an engineering leader actually needs isn't the point estimate. It's the error bars around it, and the rate at which both are moving.
What a time horizon actually measures
METR, a nonprofit AI research group, defines a model's "time horizon" as the length of task, measured in the time a skilled human would need, that the model can complete with 50% reliability. A horizon of five hours doesn't mean the model finishes every five-hour task. It means that on METR's suite of real, timed tasks, it succeeds on half of the ones sized around five hours and starts failing more often as tasks get longer. That's a cleaner metric than a leaderboard score, because it's denominated in something an engineering leader already budgets in: time.
The number under the headline, and what just changed
In the Time Horizon 1.1 update published on its blog on January 29, 2026, METR reported Claude Opus 4.5's horizon at 320 minutes, a little over five hours, with a 90% confidence interval of 170 to 729 minutes, roughly three to twelve hours. That's the figure behind most of the "AI agents can now work for hours" coverage. What the coverage tends to skip is the interval, and the fact that it just got tighter. Under METR's previous methodology (TH1), the same model's estimate was 289 minutes with an interval running from 110 to 1,268 minutes, an upper bound 4.4 times the point estimate. Under the revised methodology (TH1.1), the upper bound is 2.3 times the point estimate. The update didn't happen because the model changed. It happened because METR grew its task suite from 170 to 228 tasks, and specifically doubled its count of long tasks (8 or more hours) from 14 to 31, which is the range where the old suite was thinnest and the estimate was least trustworthy.
The growth rate matters more than the point estimate
The more decision-relevant number in the same report is how fast the horizon is moving. METR's hybrid trend across its full history shows a doubling time of 196 days, about seven months. Narrow the window and it moves faster: 131 days measured from 2023 onward under the new methodology (versus 165 days under the old one), and 89 days from 2024 onward (versus 109 days previously). Read plainly, that's an acceleration, not a steady trend, and the newer methodology finds it moving quicker than the old one did, not slower. If the 2024-onward rate holds, a five-hour horizon today implies something like a two-day horizon within a year, which is the kind of arithmetic worth sitting with rather than dismissing, and also worth not betting a roadmap on, because a trend line is not a delivery date.
Where METR's own ruler runs out
METR is explicit about the limits of its own instrument. On its time-horizons page, last updated May 8, 2026, it states that measurements above 16 hours are unreliable with its current task suite, because the suite is saturating at the long end. Even inside the 8-hour-plus band it just expanded, only 5 of its 31 long tasks have a measured human baseline; the rest are estimated. That matters because most of the higher numbers now circulating online (the ones claiming a full day or more of unsupervised work) sit above the range METR itself is willing to stand behind. The honest reading of the January 2026 report is a model that reliably handles a few hours of work, with a wide and only recently narrowed error bar, moving fast, on a scale that runs out of usable precision well before a full workday.
A benchmark is not your eval
Consider an eleven-person platform team at a Series B logistics company, deciding this sprint whether a three-day data-migration project can go to an agent instead of a person. METR's five-hour figure, on its own, answers a different question than the one that team is actually asking. It describes performance on METR's own suite of coding and reasoning tasks, evaluated inside METR's own harness, with METR's own definition of success. Nobody has run that measurement against this team's actual repository, its deployment pipeline, or the on-call rotation that has to absorb whatever the agent gets wrong. A published benchmark is a starting assumption, not a substitute for an eval harness built against your own golden set. Teams that skip that step are the ones who find out the hard way which parts of the five-hour figure were specific to METR's tasks.
What actually fits an agent, and what doesn't
The honest use of this data is narrower than the headline version. A rising time horizon is a reasonable argument for handing an agent a well-scoped, multi-hour task with a human checking the output at the end, which is exactly the shape of engagement where agents tend to hold up in production: tight scope, capped steps, a person signing off on anything irreversible. It is a poor argument for routing an entire backlog to an agent because a benchmark crossed a round number, and it says nothing at all about the parts of a senior engineer's job that were never in METR's task suite to begin with: judgment calls about what not to build, arguing a design down in a review, or knowing which two systems will fight each other in six months. Worth naming too, because a study that only ever points one direction should make you suspicious: this is a different METR study from the one that found experienced open-source developers were measurably slower with AI tools despite feeling faster. Capability on a benchmark and productivity on your actual codebase are not the same curve, and treating them as one is how a roadmap gets built on a number that was never about your team.
If your team is weighing where an agent genuinely earns its keep against where a senior engineer still has to be the one holding the pen, that's a scoping conversation worth having before the sprint plan gets written, not after it stalls. Our Silicon Valley team works through exactly that split with engineering leaders regularly, and we'll give you a straight read on the split for your roadmap within 48 hours of a call.