The Horizon
A fitted 50 percent success threshold for METR benchmark tasks, not elapsed agent runtime
What is the current Horizon reading?
METR's 50 percent time horizon is the estimated human-expert task duration at which the evaluated agent's fitted success probability is 50 percent. In April 2026, the estimate was 1044.8 hours; it is not elapsed agent runtime. The current suite primarily covers well-specified software engineering, machine learning, and cybersecurity tasks, and METR says measurements above 16 hours are unreliable.
Measurement basis: METR's 50 percent time horizon is the estimated human-expert task duration at which the evaluated agent's fitted success probability is 50 percent; it is not elapsed agent runtime. The current task suite primarily covers self-contained, well-specified software engineering, machine learning, and cybersecurity tasks with clear success criteria, and METR says measurements above 16 hours are unreliable.
METR's April 2026 estimate places the evaluated agent's 50 percent task-completion horizon at 1044.8 hours of human-expert task duration — a fitted threshold, not elapsed agent runtime.
METR defines its 50 percent time horizon as the estimated human-expert task duration at which the evaluated agent's fitted success probability is 50 percent. In April 2026, the estimate was 1044.8 hours. That number is a human task-duration threshold, not elapsed agent runtime.
The current suite primarily covers self-contained, well-specified software engineering, machine learning, and cybersecurity tasks with clear success criteria. Results should not be generalized from that benchmark scope to every kind of professional or real-world work.
METR says measurements above 16 hours are unreliable with the current task suite. The long-horizon point estimate is fitted from benchmark results; it is not evidence that an agent ran continuously for the estimated duration, and applications that require dependable completion should not treat the benchmark threshold as a guarantee.
A capability benchmark is not a labor-market outcome. The Adoption Curve reports business deployment, The AI Cut reports announced layoffs attributed to AI, and The Tech Drought reports information-sector openings. Movement across those measures does not by itself establish that benchmark capability caused an employment outcome.
Explore Further
Is this happening to you?
Are AI tools now doing tasks you used to be paid for?
How has The Horizon changed over time?
Most affected counties
Counties with the highest labor scores in the County Distress Index.
Explore all 3,144 counties →| Period | Value | YoY Change |
|---|---|---|
| Apr 2026 | 1044.8 hours | +925.1 hours |
| Feb 2026 | 718.8 hours | +658.4 hours |
| Dec 2025 | 352.2 hours | +313.4 hours |
| Nov 2025 | 293 hours | +254.2 hours |
| Aug 2025 | 203 hours | +182.7 hours |
| May 2025 | 101.2 hours | +94.2 hours |
| Apr 2025 | 119.7 hours | +116.5 hours |
| Feb 2025 | 60.4 hours | +56.8 hours |
| Dec 2024 | 38.8 hours | +34.8 hours |
| Oct 2024 | 20.5 hours | +16.5 hours |
| Sep 2024 | 20.3 hours | — |
| Jun 2024 | 11.4 hours | — |
Frequently Asked Questions
What is the AI Task Horizon?
METR's task-completion horizon is the estimated human-expert task duration at which the evaluated agent's fitted success probability is 50 percent. The April 2026 estimate is 1044.8 hours; that is a statistical task-duration threshold, not elapsed agent runtime.
Why does the task horizon matter for jobs?
The benchmark is evidence about capability on its task suite, so it can inform workforce research. It does not measure business adoption, job substitution, layoffs, or household distress, and a higher horizon does not by itself prove an employment effect.
Where does the AI Task Horizon data come from?
METR publishes benchmark results for primarily self-contained, well-specified software engineering, machine learning, and cybersecurity tasks with clear success criteria. METR warns that measurements above 16 hours are unreliable with the current task suite.
Quick poll
Is this affecting you or your household?
Discussion
Get the numbers when they move.
New data drops, indicator updates, and ADI score changes — delivered when it matters. No spam.
or Create an Account for full access
Loading comments…