Epoch estimates frontier AI benchmark progress nearly doubled in pace around April 2024
Epoch AI’s December 2025 retrospective estimated that the frontier of its composite Epoch Capabilities Index improved 1.85 times faster after a best-fit breakpoint on 8 April 2024: 8.337 ECI points per year before the breakpoint and 15.459 afterward.
The Reality Check
Context
The Epoch Capabilities Index combines results from many benchmarks using a logistic latent-variable model that estimates model capability, benchmark difficulty, and benchmark discrimination. Epoch states that the resulting scale is arbitrary and currently anchored to Claude 3.5 Sonnet at 130 and GPT-5 at 150.
For the December 2025 retrospective, Epoch refit the index using 149 models released from December 2021 through December 2025, selected 17 models that established new running maxima, and tested 5,000 candidate breakpoint dates. The minimum-residual continuous two-segment fit placed the breakpoint on 8 April 2024, with slopes of 8.337 and 15.459 ECI points per year. The ratio was 1.85×. Across 2,000 resampled frontier datasets, the two-segment model beat a single-line model on AIC in 90% of runs and on BIC in approximately 80%. The 90% intervals were 1.3–3.1× for the ratio and 17 March 2023–19 August 2024 for the breakpoint.
Epoch’s April 2026 follow-up tested four metrics, eight trend families, and four data-preparation modes. ECI, log METR time horizon, and the mathematics index showed acceleration relative to a global linear trend, while WeirdML V2 did not. Reasoning and non-reasoning models were best represented by separate trends in the positive metrics, but several superlinear forms also fit well.
The result is therefore sensitive to benchmark composition, model inclusion, elicitation and scoring choices, refitting, and the selected trend family. No independent reproduction of Epoch’s exact public ECI series, April 2024 breakpoint, and 1.85× slope ratio was located.
THE TAKEAWAY
The evidence supports a faster frontier trend in selected, benchmark-heavy capabilities after approximately 2024 than one constant linear rate across the full 2021–2025 period. It does not show that every AI capability accelerated by 85%, that progress changed suddenly on precisely 8 April 2024, or that reasoning models or reinforcement learning caused the change.
The strongest supported conclusion is narrower: several readily verifiable mathematics, programming, and related capability measures show evidence of accelerated measured progress, while other measures do not yet show the same pattern.
Continue the Thread
Frontier AI Capability Progress Over TimeTracks evidence about how quickly and consistently frontier AI capabilities are changing over time, including acceleration, slowdowns, plateaus, and breaks in long-run trends.
Sources
AI capabilities progress has sped up
Epoch AI
Used for: Disclosure date, retrospective dataset scope, frontier-point selection, breakpoint-search procedure, corrected 8 April 2024 estimate, pre- and post-breakpoint slopes, 1.85× ratio, model comparisons, resampling results, bootstrap intervals, assumptions, limitations, and erratum.
Have AI Capabilities Accelerated?
Epoch AI
Used for: Later testing across ECI, METR time horizon, a combined mathematics index, and WeirdML V2; comparison of eight trend families and four dataset-preparation modes; reasoning-versus-non-reasoning results; domain-generalization limits; and the non-accelerating WeirdML result.
ECI: Epoch Capabilities Index
Epoch AI / GitHub
Used for: Public implementation of the ECI model, fitting scripts, benchmark-data loading, capability, difficulty and discrimination parameters, bootstrap outputs, confidence intervals, and scale anchoring.
A Rosetta Stone for AI Benchmarks
arXiv / Epoch AI and collaborating researchers
Used for: Technical basis for placing heterogeneous benchmarks on one latent capability scale, benchmark-stitching assumptions, simulation-based validation, applications to capability trends, and limitations of representing model capability with one number.
ECI Documentation – Methodology
Epoch AI
Used for: Current logistic model specification, nonlinear least-squares fitting, regularization, arbitrary scale construction, anchor models, and bootstrap confidence intervals.
Epoch Capabilities Index
Epoch AI
Used for: Purpose of the composite index, current benchmark breadth, interpretation of absolute scores and slopes, ongoing rescaling, domain-specific limitations, and the public implementation link.
Measuring AI Ability to Complete Long Software Tasks
NeurIPS 2025
Used for: Separate evidence that software-agent task horizons followed a rapid historical growth trend and may have accelerated during 2024; used as convergent context based on a different metric, not as a reproduction of ECI.
System Card: Claude Mythos Preview
Anthropic
Used for: A separate organization’s internal ECI-style trend analysis, its 1.86× to 4.3× slope-ratio range across breakpoint choices, and sensitivity to benchmark selection; not used as a reproduction of Epoch’s exact result.
Last checked Methodology 2.0.0