Skip to content
SingularityOBSERVATORY
Menu
TRACKER/AGENT-HORIZON/RELIABILITY

DECISION VIEW / AGENTS

Horizon depends on the reliability requirement

The same frontier cohort looks radically different at 50 and 80 percent success. Suite coverage imposes a firm boundary on interpretation.

Task duration & success rate

Human-expert work time, not how long an agent runs.

30 min2 h8 h32 h
Estimate & source interval16 h reliability boundaryLogarithmic scale

16h · UNRELIABLE ABOVE THIS. Frontier is a reported cohort; named models are separate measurements, not a ranking.

Analysis summary Assessed 2026-05-08
OBSERVATION

approx. 1.5 h

Public frontier at 80% success

The stricter measure is often more relevant to decisions than P50.

Choose the reliability level before interpreting autonomy, and stop where the benchmark's long tasks no longer support the estimate.

PUBLIC EVIDENCE STATUSDEVELOPING

4 source-bound P50/P80 estimates with intervals under the same TH1.1 contract; suite coverage and the 16-hour boundary still limit generalisability.

Claim evidence · agent-horizon/claim/evidence-state ↓
LATEST OBSERVATION
2026-05-08
LAST SOURCE CHECK
2026-08-23
CADENCE
monthly
COMPATIBLE OBSERVATIONS
4non-contextual P50/P80 estimates under the same TH1.1 contract
TREND STATUS
NOT ESTABLISHED4 estimates are comparable across models and reliability levels, but are not a time series; Opus 4.5 is shown only as context.
Download machine-readable evidence ↓
RELIABILITY DROP
approx. 12 h @ 50%→approx. 1.5 h @ 80%
8× shorter

Point estimates for the same public frontier cohort.

MEASUREMENT DEVELOPMENT
Opus 4.5: 320 min→Frontier cohort: approx. 720 min
different assessments

Shown as context, not as a pure model trend.

50% SUCCESSSIGNAL

How far at coin-flip reliability?

The public frontier is around 12 hours, but the interval is 5–61 hours and extends into the saturation region.

OBSERVATIONapprox. 12 h [5–61]

Public frontier, 50% · Cohort result, wide uncertainty.

80% SUCCESSSIGNAL

How far at higher reliability?

The point estimate falls to approximately 1.5 hours, with an interval of 50 minutes–2 hours 40 minutes.

OBSERVATIONapprox. 1.5 h [50m–2h40m]

Public frontier, 80% · Stricter success requirement.

SENSITIVITYLIMITATION

How much does the success requirement matter?

The rounded P50 and P80 point estimates differ by a factor of eight.

INTERPRETATION8×

P50/P80 point estimate · 720 / 90 expert-minutes.

SUITE COVERAGELIMITATION

How much long-duration work is in the measurement basis?

31 of 228 tasks take 8+ hours; only five tasks in the entire suite are long tasks with a human baseline.

INTERPRETATION228 tasks

TH1.1 task suite · 31 take 8+ hours; five have a human baseline.

INTERPRETATION13.6% / 2.2%

Long-task coverage · Share taking 8h+ / share taking 8h+ with a human baseline.

MEASUREMENT GAPMEASUREMENT GAP

When should the figure not be used?

METR marks horizons above 16 hours as unreliable with the current suite.

MEASUREMENT GAP>16 h unreliable

Measurement boundary · Source-stated limitation.

SOURCES / PROVENANCE4 source records
primary-measurementMETRTime Horizon 1.1
Retrieved
2026-08-12
Data
2026-01-29
Licence
Cited public research release; analysis repository license applies to its code/data

Public, task-suite-specific agent horizon estimates; this module records only stated values, not a copied raw dataset.

Open original source ↗
measurement-methodologyMETRTask-Completion Time Horizons of Frontier AI Models
Retrieved
2026-08-12
Data
2026-05-08
Licence
Cited public web source; repository license applies separately

Defines time horizon and explicitly marks estimates above 16 hours unreliable in the current suite.

Open original source ↗
primary-measurementMETRFrontier Risk Report: February to March 2026
Retrieved
2026-08-12
Data
2026-03-31
Licence
Cited public research report

Cohort-level public-frontier TH1.1 results at 50% and 80% reliability; saturation and wide intervals are material limitations.

Open original source ↗
reviewed-research-sourceMETRMETR — Time Horizon 1.1 benchmark results
Retrieved
2026-08-23
Data
2026-05-08
Licence
citation-only-review-required

METR's official machine-readable TH1.1 result artifact reports GPT-5.4 at 341.735276 expert-minutes P50 [186.581591–768.779526] and 53.877851 expert-minutes P80 [23.957027–108.679232]. Both estimates remain below METR's 16-hour reliability boundary and fit the accepted Agent Horizon adapter without a methodology change.

Open original source ↗
VERIFIABLE EVIDENCE / 10 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Observation

METR TH1.1 reports 320 minutes [170–729] for Claude Opus 4.5 at a modelled 50% success rate.

METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • A specific model, suite and agent setup.
agent-horizon/claim/opus-p50
Observation

METR's February–March 2026 assessment places the public frontier at approximately 12 hours [5–61] at 50% success.

  • A cohort estimate with a wide interval and a suite nearing saturation.
agent-horizon/claim/public-frontier-p50
Observation

The same METR assessment places the public frontier at approximately 1.5 hours [50 minutes–2 hours 40 minutes] at 80% success.

  • A cohort estimate, not a ranking of named models.
agent-horizon/claim/public-frontier-p80
Calculation

The public frontier point estimate is eight times longer at 50% than at 80% success, showing strong sensitivity to the reliability requirement.

CALCULATION12 hours / 1.5 hours = 8
  • The ratio uses rounded point estimates; intervals are wide.
agent-horizon/claim/reliability-gap
Method

TH1.1 has 228 tasks, of which 31 are estimated to take at least eight hours and five of these have a human baseline.

METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • Not a representative distribution of labour-market tasks.
agent-horizon/claim/suite-coverage
Calculation

Only 13.6% of TH1.1 tasks take eight hours or longer, and 2.2% of the entire suite consists of such tasks with a human baseline.

CALCULATION31 / 228 = 13.6%; 5 / 228 = 2.2%
METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • Task counts say nothing about representativeness or quality.
agent-horizon/claim/long-task-coverage
Method

Agent Horizon shows 4 compatible P50/P80 estimates across 2 public model cohorts under the same TH1.1 contract. They are concurrent model and reliability observations, not a time series; suite coverage and the 16-hour boundary limit generalisability.

  • P50 and P80 are reliability levels, not separate points in time.
  • The suite is limited to software, ML and cybersecurity tasks.
agent-horizon/claim/evidence-state
Observation

METR TH1.1 reports 341.735276 expert-minutes [186.581591–768.779526] for GPT-5.4 at 50% success.

  • Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p50
Observation

METR TH1.1 reports 53.877851 expert-minutes [23.957027–108.679232] for GPT-5.4 at 80% success.

  • Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p80
TRACEABLE FACT LOG / 7 SOURCE FACTSOpen log +
agent-horizon/fact/metr/opus-4-5-p50
320 [170,729] minutes

Claude Opus 4.5 TH1.1 P50

METR · Changes to Model Horizon Estimates table
agent-horizon/fact/metr/suite
228 tasks; 31 at 8h+; 5 human baselined

TH1.1 task suite

METR · Task suite description
agent-horizon/fact/metr/public-frontier-p50
about 12h [5h,61h]

Public frontier TH1.1 P50

METR · Table 1: public frontier Feb–Mar 2026
agent-horizon/fact/metr/public-frontier-p80
about 1.5h [50m,2h40m]

Public frontier TH1.1 P80

METR · Table 1: public frontier Feb–Mar 2026
agent-horizon/fact/metr/limit
above 16h unreliable

Long task reliability boundary

METR · Methodological Details
agent-horizon/fact/metr/gpt-5-4-p50
341.735276 [186.581591,768.779526] minutes

GPT-5.4 TH1.1 P50

METR · results.gpt_5_4.metrics.p50_horizon_length and results.gpt_5_4.metrics.p80_horizon_length
agent-horizon/fact/metr/gpt-5-4-p80
53.877851 [23.957027,108.679232] minutes

GPT-5.4 TH1.1 P80

METR · results.gpt_5_4.metrics.p50_horizon_length and results.gpt_5_4.metrics.p80_horizon_length
INTERPRETATION LIMIT

The frontier cohort is not a ranking of named models.

Intervals are wide and saturation affects P50.

Software tasks cannot be generalised directly to ordinary knowledge work.

MODULE STATUSSPARSE

Public frontier: approx. 12 h at 50% success, but approx. 1.5 h at 80%

VARDARK OBSERVATORY / AGENT-HORIZON / RELIABILITY