The public agent horizon has lengthened
METR's February–March 2026 assessment places the public frontier at approximately 12 hours at 50 percent success, but only approximately 1.5 hours at 80 percent success.
Inspect claim and source →CLAIM OBSERVATORY · REVIEWED ASSESSMENT
Accepted evidence shows growth in estimated accelerator ownership and demonstrates bounded agent-task capability. These measurements do not establish highly reliable autonomy or an independently validated frontier AI-research loop. AI 2027 therefore remains a scenario to watch, while slower and discontinuous pathways remain plausible alternatives.
This assessment was reviewed on the date shown. Earlier wording remains in the history below; newer data does not silently renew the judgment.
METR's February–March 2026 assessment places the public frontier at approximately 12 hours at 50 percent success, but only approximately 1.5 hours at 80 percent success.
Inspect claim and source →The Epoch AI series gives a median of 20.9 million H100e at the end of 2025, 208 percent above the same quarter a year earlier. H100e is a comparison model, not measured useful work.
Inspect claim and source →The point estimate is eight times longer at 50 than at 80 percent success, and METR marks estimates above 16 hours as unreliable with the current task set.
Inspect claim and source →The three benchmarks cannot be combined into a single automation share, and current results do not document a robust, independent research loop.
Inspect claim and source →These charts share a reading page, not a numeric axis or causal conversion. The Compute dates describe estimated ownership; the Agent Horizon points describe one task suite at two success thresholds.
Global ownership estimate · H100e comparison model
H100e is a normalization, not a count of H100 cards or verified installed, operational or useful compute. The displayed range is model uncertainty.
Compute method and full series →METR task suite · expert task duration · same cohort
The ranges are broad; expert task duration is not agent runtime. These simultaneous thresholds are not a time trend. Estimates above 16 expert-hours are unreliable with the current suite.
Reliability and suite limits →The Transition Map records the bridge as qualified relationships, not a conversion from hardware to autonomous work.
More deployable hardware is not converted into model capability by any accepted Observatory transfer function.
Training allocation, algorithms, data, model architecture and inference use are absent.Agent task horizon supplies one operational capability indicator under a specific suite and reliability level; it is not general capability.
The estimates are suite- and agent-scaffold-specific.No accepted relationship converts general task horizon into end-to-end frontier research automation.
The benchmark families use different tasks, evaluators and success criteria.AI 2027 includes a research feedback mechanism. Its current assessment was reviewed : Accepted evidence shows growth in estimated accelerator ownership and demonstrates bounded agent-task capability. These measurements do not establish highly reliable autonomy or an independently validated frontier AI-research loop. AI 2027 therefore remains a scenario to watch, while slower and discontinuous pathways remain plausible alternatives.
Evidence that fits part of a pathway does not confirm the full pathway. Slow or uneven development keeps slower or uneven development visible as an alternative.
Include a dated, source-linked benchmark or observation only when its task and success definition are stated; source plans remain plans. Relevance means bearing on a named assumption, not confirming the full scenario.
Strengthen this mechanism only if reproducible evidence shows an agent carrying a substantial frontier-research task from hypothesis to independently validated result at high reliability, with task and scaffold scope explicit. Weaken the mechanism only if comparable, high-quality evaluations show persistent failure at the reliability levels that research tasks require, persistent failure to validate or reproduce research, or stagnation or regression in task reach at stricter success thresholds. A persistent gap between the 50% and 80% success horizons does not by itself warrant weakening, because both horizons can grow while that gap remains. A larger ownership estimate or facility plan alone does not settle the question.
These are checks for the next review, not forecast dates or claims that the observations have happened.
Compared with the latest retained knowledge date, 26 Sept 2026, 01:40 UTC. Measurement periods can be older than the date a source entered the Observatory. Your selected date is included in the page link.
No new retained observations for this claim after the selected date.
No source correction or coverage change after the selected date.
Retained scenario assessment history: Baseline (22 Aug 2026): Agent progress and compute scaling fit parts of the pathway, but the decisive milestone of robust automation of frontier AI research has not been publicly confirmed. Revision (26 Sept 2026, 01:40 UTC): Accepted evidence shows growth in estimated accelerator ownership and demonstrates bounded agent-task capability. These measurements do not establish highly reliable autonomy or an independently validated frontier AI-research loop. AI 2027 therefore remains a scenario to watch, while slower and discontinuous pathways remain plausible alternatives.