METR finds frontier AI task-completion horizons doubling about every seven months
METR’s human-time-calibrated benchmark found that the frontier 50% task-completion horizon for model agents on its software and research tasks roughly doubled every seven months from 2019 to early 2025.
Claim: ConfirmedEvidence: Strong