DECISION VIEW / AGENTS
Horizon depends on the reliability requirement
The same frontier cohort looks radically different at 50 and 80 percent success. Suite coverage imposes a firm boundary on interpretation.
Task duration & success rate
Human-expert work time, not how long an agent runs.
Frontier · 50% success
12 h (5 h–61 h)Frontier · 80% success
1.5 h (50 min–2.7 h)GPT-5.4 · 50% success
5.7 h (3.1 h–12.8 h)GPT-5.4 · 80% success
53.9 min (24 min–1.8 h)16h · UNRELIABLE ABOVE THIS. Frontier is a reported cohort; named models are separate measurements, not a ranking.
Analysis summary Assessed 2026-05-08
approx. 1.5 h
Public frontier at 80% successThe stricter measure is often more relevant to decisions than P50.
Choose the reliability level before interpreting autonomy, and stop where the benchmark's long tasks no longer support the estimate.
4 source-bound P50/P80 estimates with intervals under the same TH1.1 contract; suite coverage and the 16-hour boundary still limit generalisability.
Claim evidence · agent-horizon/claim/evidence-state ↓- LATEST OBSERVATION
- 2026-05-08
- LAST SOURCE CHECK
- 2026-08-23
- CADENCE
- monthly
- COMPATIBLE OBSERVATIONS
- 4non-contextual P50/P80 estimates under the same TH1.1 contract
- TREND STATUS
- NOT ESTABLISHED4 estimates are comparable across models and reliability levels, but are not a time series; Opus 4.5 is shown only as context.
- MEASUREMENT GAP
- 2tasks longer than 16 hourseconomic generalisability
Point estimates for the same public frontier cohort.
Shown as context, not as a pure model trend.
How far at coin-flip reliability?
The public frontier is around 12 hours, but the interval is 5–61 hours and extends into the saturation region.
Public frontier, 50% · Cohort result, wide uncertainty.
How far at higher reliability?
The point estimate falls to approximately 1.5 hours, with an interval of 50 minutes–2 hours 40 minutes.
Public frontier, 80% · Stricter success requirement.
How much does the success requirement matter?
The rounded P50 and P80 point estimates differ by a factor of eight.
P50/P80 point estimate · 720 / 90 expert-minutes.
How much long-duration work is in the measurement basis?
31 of 228 tasks take 8+ hours; only five tasks in the entire suite are long tasks with a human baseline.
TH1.1 task suite · 31 take 8+ hours; five have a human baseline.
Long-task coverage · Share taking 8h+ / share taking 8h+ with a human baseline.
When should the figure not be used?
METR marks horizons above 16 hours as unreliable with the current suite.
Measurement boundary · Source-stated limitation.
- Retrieved
- 2026-08-12
- Data
- 2026-01-29
- Licence
- Cited public research release; analysis repository license applies to its code/data
Public, task-suite-specific agent horizon estimates; this module records only stated values, not a copied raw dataset.
Open original source ↗- Retrieved
- 2026-08-12
- Data
- 2026-05-08
- Licence
- Cited public web source; repository license applies separately
Defines time horizon and explicitly marks estimates above 16 hours unreliable in the current suite.
Open original source ↗- Retrieved
- 2026-08-12
- Data
- 2026-03-31
- Licence
- Cited public research report
Cohort-level public-frontier TH1.1 results at 50% and 80% reliability; saturation and wide intervals are material limitations.
Open original source ↗- Retrieved
- 2026-08-23
- Data
- 2026-05-08
- Licence
- citation-only-review-required
METR's official machine-readable TH1.1 result artifact reports GPT-5.4 at 341.735276 expert-minutes P50 [186.581591–768.779526] and 53.877851 expert-minutes P80 [23.957027–108.679232]. Both estimates remain below METR's 16-hour reliability boundary and fit the accepted Agent Horizon adapter without a methodology change.
Open original source ↗VERIFIABLE EVIDENCE / 10 CLAIMSOpen evidence +
Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.
METR TH1.1 reports 320 minutes [170–729] for Claude Opus 4.5 at a modelled 50% success rate.
- A specific model, suite and agent setup.
agent-horizon/claim/opus-p50METR's February–March 2026 assessment places the public frontier at approximately 12 hours [5–61] at 50% success.
- A cohort estimate with a wide interval and a suite nearing saturation.
agent-horizon/claim/public-frontier-p50The same METR assessment places the public frontier at approximately 1.5 hours [50 minutes–2 hours 40 minutes] at 80% success.
- A cohort estimate, not a ranking of named models.
agent-horizon/claim/public-frontier-p80The public frontier point estimate is eight times longer at 50% than at 80% success, showing strong sensitivity to the reliability requirement.
12 hours / 1.5 hours = 8- The ratio uses rounded point estimates; intervals are wide.
agent-horizon/claim/reliability-gapTH1.1 has 228 tasks, of which 31 are estimated to take at least eight hours and five of these have a human baseline.
- Not a representative distribution of labour-market tasks.
agent-horizon/claim/suite-coverageOnly 13.6% of TH1.1 tasks take eight hours or longer, and 2.2% of the entire suite consists of such tasks with a human baseline.
31 / 228 = 13.6%; 5 / 228 = 2.2%- Task counts say nothing about representativeness or quality.
agent-horizon/claim/long-task-coverageMETR marks estimates above 16 hours as unreliable with the current task suite.
- Not a statement about actual maximum autonomy.
agent-horizon/claim/long-task-limitAgent Horizon shows 4 compatible P50/P80 estimates across 2 public model cohorts under the same TH1.1 contract. They are concurrent model and reliability observations, not a time series; suite coverage and the 16-hour boundary limit generalisability.
- P50 and P80 are reliability levels, not separate points in time.
- The suite is limited to software, ML and cybersecurity tasks.
agent-horizon/claim/evidence-stateMETR TH1.1 reports 341.735276 expert-minutes [186.581591–768.779526] for GPT-5.4 at 50% success.
- Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p50METR TH1.1 reports 53.877851 expert-minutes [23.957027–108.679232] for GPT-5.4 at 80% success.
- Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p80TRACEABLE FACT LOG / 7 SOURCE FACTSOpen log +
agent-horizon/fact/metr/opus-4-5-p50Claude Opus 4.5 TH1.1 P50
METR · Changes to Model Horizon Estimates tableagent-horizon/fact/metr/suiteTH1.1 task suite
METR · Task suite descriptionagent-horizon/fact/metr/public-frontier-p50Public frontier TH1.1 P50
METR · Table 1: public frontier Feb–Mar 2026agent-horizon/fact/metr/public-frontier-p80Public frontier TH1.1 P80
METR · Table 1: public frontier Feb–Mar 2026agent-horizon/fact/metr/limitLong task reliability boundary
METR · Methodological Detailsagent-horizon/fact/metr/gpt-5-4-p50GPT-5.4 TH1.1 P50
METR · results.gpt_5_4.metrics.p50_horizon_length and results.gpt_5_4.metrics.p80_horizon_lengthagent-horizon/fact/metr/gpt-5-4-p80GPT-5.4 TH1.1 P80
METR · results.gpt_5_4.metrics.p50_horizon_length and results.gpt_5_4.metrics.p80_horizon_lengthThe frontier cohort is not a ranking of named models.
Intervals are wide and saturation affects P50.
Software tasks cannot be generalised directly to ordinary knowledge work.
Public frontier: approx. 12 h at 50% success, but approx. 1.5 h at 80%