METR finds frontier AI task-completion horizons doubling about every seven months
METR’s human-time-calibrated benchmark found that the frontier 50% task-completion horizon for model agents on its software and research tasks roughly doubled every seven months from 2019 to early 2025.
The Reality Check
Context
METR’s original study combined 170 software and research tasks and modelled agent success against the amount of time an expert human would need to complete each task. The resulting 50% horizon provided a common human-time scale for comparing model-agent systems across generations.
The peer-reviewed NeurIPS version retained the long-run trend and reported an approximately 110-minute horizon for o3. METR’s TH1.1 update expanded the suite to 228 tasks; its stitched long-run estimate was about 196.5 days per doubling, close to the original 195.8-day estimate. Faster post-2023 and post-2024 slopes were more sensitive to task composition, model selection, and modelling assumptions.
A time horizon is a human-time difficulty measure, not the amount of wall-clock time an agent can work without supervision.
The benchmark mostly uses self-contained, automatically gradable software and research tasks. Results depend on the model, scaffold, tools, prompts, and resource limits. Human baselines are also noisy; in TH1.1, only five of the 31 tasks estimated at eight hours or more had measured human baselines. BRIDGE recovered a similar exponential trend with a different statistical model, but its calibration still partly uses METR’s human-time annotations.
THE TAKEAWAY
The durable result is rapid improvement on METR-like software and research tasks across multiple model generations. It does not establish that a one-hour horizon equals one hour of dependable workplace autonomy, that every one-hour professional task is within reach, or that the faster recent slope will continue indefinitely.
Continue the Thread
AI Agents in Software EngineeringTracks the capability of AI agents to complete software-engineering tasks across increasing scope, duration, and autonomy.
Sources
Measuring AI Ability to Complete Long Software Tasks
NeurIPS 2025
Used for: Peer-reviewed methodology, composition of the original task suite, model results, long-run doubling trend, robustness work, and limitations.
Time Horizon 1.1
METR
Used for: Expansion from 170 to 228 tasks, revised model estimates, long-run and recent-window doubling rates, task-composition sensitivity, and updated Claude 3.7 Sonnet estimates.
BRIDGE: Predicting Human Task Completion Time From Model Performance
ICML 2026 / arXiv
Used for: Separate-team reproduction of the exponential task-horizon trend using a different item-response framework and additional benchmark data; also used to assess how independent the replication is from METR’s calibration data.
METR Time Horizon Analysis
METR / GitHub
Used for: Public analysis code, data pipeline, logistic fitting procedure, bootstrap analysis, and reproducibility of the reported figures and time-horizon estimates.
Clarifying Limitations of Time Horizon
METR
Used for: Clarification that time horizon is not wall-clock autonomous runtime, domain and task-distribution limits, uncertainty in individual estimates, reliability limits, and warnings against treating the metric as a direct automation rate.
Impact of Modelling Assumptions on Time Horizon Results
METR
Used for: The corrected regularization mistake, benchmark-saturation concerns, sensitivity to curve fitting and task-length noise, and the scale of uncertainty in recent model estimates.
Measuring AI Ability to Complete Long Tasks
METR
Used for: Original March 19, 2025 disclosure, public explanation of the 50% task-completion time horizon, original Claude 3.7 Sonnet result, and the approximately seven-month trend claim.
Last checked Methodology 2.0.0