Skip to content
SingularityOBSERVATORY
Menu

Tracker

Measurements & trends

6 active tracks
Visible tracks

AI compute capacity

Estimated ownership · Q4 2025
Observed estimate · Q4 202541.4 ZOPSWorld
5th–95th percentile32.1–53.8 ZOPSView evidence

+208% global annual growthSame quarter, previous year

014.829.644.459.2Q1 2022202320242025Q4 2025US-based cloud groups · observed estimatesChina · observed estimatesOther / unallocated · observed estimatesObserved estimate · World · 2022-03-31 · 0.12 ZOPSObserved estimate · World · 2022-06-30 · 0.26 ZOPSObserved estimate · World · 2022-09-30 · 0.4 ZOPSObserved estimate · World · 2022-12-31 · 0.55 ZOPSObserved estimate · World · 2023-03-31 · 0.81 ZOPSObserved estimate · World · 2023-06-30 · 1.3 ZOPSObserved estimate · World · 2023-09-30 · 2.1 ZOPSObserved estimate · World · 2023-12-31 · 3.1 ZOPSObserved estimate · World · 2024-03-31 · 4.8 ZOPSObserved estimate · World · 2024-06-30 · 6.9 ZOPSObserved estimate · World · 2024-09-30 · 9.8 ZOPSObserved estimate · World · 2024-12-31 · 13.4 ZOPSObserved estimate · World · 2025-03-31 · 18.6 ZOPSObserved estimate · World · 2025-06-30 · 24.5 ZOPSObserved estimate · World · 2025-09-30 · 32.5 ZOPSObserved estimate · World · 2025-12-31 · 41.4 ZOPS
Q4 2025
Link to selected viewUpdating view · data through 2025-12-31. New data or corrections may change this view.
View chart values (16)

Values for World through 2025-12-31. Plans and scenarios are separate entries, never added to observed capacity. Select a date to inspect and share that point.

World · ZOPS · with uncertainty bounds
Date / inspectType / entryCapacity · ZOPSBoundsEvidence
Observed estimateWorld0.1240.073–0.2125th–95th percentileEvidence
Observed estimateWorld0.2570.153–0.4365th–95th percentileEvidence
Observed estimateWorld0.40.246–0.6545th–95th percentileEvidence
Observed estimateWorld0.550.358–0.8585th–95th percentileEvidence
Observed estimateWorld0.810.56–1.185th–95th percentileEvidence
Observed estimateWorld1.290.934–1.795th–95th percentileEvidence
Observed estimateWorld2.061.53–2.85th–95th percentileEvidence
Observed estimateWorld3.122.33–4.145th–95th percentileEvidence
Observed estimateWorld4.793.58–6.455th–95th percentileEvidence
Observed estimateWorld6.895.17–9.285th–95th percentileEvidence
Observed estimateWorld9.757.4–13.025th–95th percentileEvidence
Observed estimateWorld13.4410.19–17.935th–95th percentileEvidence
Observed estimateWorld18.6414.23–24.675th–95th percentileEvidence
Observed estimateWorld24.4718.7–32.395th–95th percentileEvidence
Observed estimateWorld32.5325.08–42.515th–95th percentileEvidence
Observed estimateWorld41.4132.1–53.795th–95th percentileEvidence

All ownership categories in the Epoch AI dataset.

Theoretical dense 8-bit peak capacity · ZOPS: 10²¹ operations per second, not effective model capability. Ownership is not physical location. ≈ 20.9M H100e at the latest global median.

Global ownership context

Q4 2025
Annual growth+208%Same quarter, previous year · Evidence
US-based cloud groups75%Global ownership · Evidence
Nominal accelerator power13.4 GWModelled, not measured electricity · Evidence
Q4 2025
Owner groupGlobal shareCapacity · ZOPSUncertainty range
US-based cloud groups ↗
75.2%
31.225–38.7
China
8.7%
3.62.7–5.7
Other / unallocated
16.1%
6.74.4–9.4
Separate measure · EU-located facilities0.06 ZOPS · documented minimum ↗
Scenario comparisons 2
SCENARIO LAYER / CROSS-MODULEDated references

Shows where external scenarios sit in time and capacity. Context is not the same as a proven threshold.

Evidence ↗
Evidence ↗
Evidence status, measurement limits & missing data
PUBLIC EVIDENCE STATUSDEVELOPING

16 compatible quarters with source archives and uncertainty intervals, but the series estimates ownership and ends at 2025 Q4.

Claim evidence · compute/claim/evidence-state ↓
LATEST OBSERVATION
2025-12-31
LAST SOURCE CHECK
2026-09-08
CADENCE
weekly
COMPATIBLE OBSERVATIONS
16quarters in the global ownership series
TREND STATUS
VALID WITHIN THIS SERIESThe quarters use the same Epoch AI basis and are comparable as estimated ownership holdings; the EU minimum and facility plans are excluded from the trend.
Download machine-readable evidence ↓
  • Measures estimated ownership, not use, availability or physical location.
  • H100e normalises peak 8-bit performance and is not the same as effective training performance.
  • Summed 5th and 95th percentiles are display bounds, not a joint probability model.
  • US cloud is a company category; other capacity cannot be reliably allocated geographically.
  • The EU series is a location-based documented minimum from a selected frontier registry, not total EU capacity or an ownership share.
  • The dashed future curve mechanically extends the last eight observed quarters; it is not a forecast.
  • Planned data centres are individual facilities and cannot be added to ownership holdings without risking double counting.
VERIFIABLE EVIDENCE / 15 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Method

The curve estimates cumulative AI accelerator ownership in the Epoch AI dataset; it does not measure actual use, access, physical location or effective model capacity.

  • Coverage follows Epoch AI's ownership model and is not a complete inventory of all computing worldwide.
compute/claim/measurement-scope
Method

H100e is a comparison model that normalises different AI chips against the Nvidia H100's theoretical maximum dense 8-bit performance; it is not a count of H100 cards.

CALCULATIONH100e = peak dense 8-bit OPS/s ÷ 1.979e15; ZOPS = H100e × 1.979e15 ÷ 1e21
  • Memory, networking, software, precision, utilisation and workload can make actual performance higher or lower than this comparison.
compute/claim/h100e-model
Method

The ownership-based series aggregate Epoch AI's global ownership categories: US cloud comprises seven named companies, China combines official and estimated smuggled categories, and Other is the global Other category.

  • Ownership categories do not indicate geographical location, access or state control.
compute/claim/series-definitions
Method

The EU series sums positive H100e capacity at facilities in the EU's 27 member states included in Epoch AI's selected frontier data centre registry.

  • The registry is not a complete inventory of all data centres or all AI capacity in the EU.
  • The series uses physical location, while the main series use ownership.
  • Non-frontier facilities use Epoch's capacity factor 1.5 and ±6 months for estimated operational dates (80% coverage); this is not the ownership curve's 5th–95th percentiles. Observation dates are not operational-date estimates.
compute/claim/eu-minimum-definition
Method

The series' 5th and 95th percentiles sum the source groups' respective bounds as a transparent display approximation.

CALCULATIONseries_low = Σ owner_low; series_high = Σ owner_high
  • The summed bounds are not a joint probabilistic confidence model.
compute/claim/uncertainty-model
Calculation

The world series median increased by 208% from the same quarter a year earlier to 2025-12-31.

CALCULATIONgrowth_pct = (current_median ÷ prior_year_median − 1) × 100
compute/claim/year-growth-latest
Calculation

The seven named US-based cloud and model groups account for 75% of the world series' estimated ownership at 2025-12-31.

CALCULATIONshare_pct = us_cloud_median ÷ world_median × 100
  • US cloud is a company category, not a claim about physical location or state control.
compute/claim/us-cloud-share-latest
Observation

31,834 H100e is the latest documented positive minimum at EU-located frontier facilities as of 2026-08-16, based on 1 facility in the archived registry.

CALCULATIONEU_documented_minimum = Σ positive listed facility H100e in EU member states
  • This is a documented minimum in a frontier registry, not an estimate of all AI compute in the EU.
  • The series is location-based and cannot be subtracted from Epoch AI's ownership-based Other category.
  • Future facility points are plans with uncertain dates and capacity, not observed measurements.
compute/claim/eu-minimum-latest
Calculation

Nominal accelerator power in the world series is 13.4 GW at 2025-12-31.

CALCULATIONaccelerator_power_GW = summed_accelerator_power_MW ÷ 1000
  • This is nominal accelerator power, not actual electricity consumption or total data centre load.
compute/claim/power-proxy-latest
Calculation

The dashed curve mechanically extends the CAGR of the last eight complete observations quarter by quarter; it is not a forecast.

CALCULATIONfuture = latest_median × annual_growth_multiple ^ elapsed_years
  • The extension assumes an unchanged historical growth rate and does not model power, capital, supply or saturation.
compute/claim/growth-scenario
Calculation

AI 2027 places the superhuman-coder milestone in March 2027 and specifies 60 million H100e globally during the Agent-3 period beginning that month; this is 2.9× the latest observed world median, but is scenario context rather than a causal takeoff threshold.

CALCULATIONdistance_multiple = AI_2027_context_H100e ÷ latest_observed_H100e
  • AI 2027 and Epoch AI use H100e as a comparison unit, but their normalisation methods are not documented as identical.
  • The alignment of date and capacity provides context; it does not mean that 60M H100e triggers the milestone.
compute/claim/ai-2027-superhuman-coder-context
Calculation

AI 2027 estimates 100 million globally available H100e at the end of 2027, 4.8× the latest observed world median in this tracker.

CALCULATIONdistance_multiple = AI_2027_forecast_H100e ÷ latest_observed_H100e
  • This is an external forecast, not a Vardark measurement.
  • AI 2027 and Epoch AI use H100e as a comparison unit, but their normalisation methods are not documented as identical.
compute/claim/ai-2027-global-compute
Method

Compute has 16 compatible quarters in the global ownership-based series from 2022-03-31 to 2025-12-31. The trend is valid within this series; the EU minimum and facility plans use a different measurement basis and are excluded.

  • Maturity describes the evidence contract, not the probability of a singularity.
  • A new source edition may require compatibility review before extending the series.
compute/claim/evidence-state
MEASUREMENT LOG / 66 TRACEABLE DATA POINTSOpen log +

Energy & Grid Pressure

Data through · 2025-12-31

Data-centre electricity demand

TWh / year · worldwide
2025 estimate485 TWh
2030 source scenario950 TWhIEA 2026 update
Historical estimatesSource scenario
20%
of planned capacity may be delayed by grid risk

Separate grid-delay risk estimate; not subtracted from electricity demand.

Forecast

All data centres, not AI alone. Only 2 historical observations; dashed lines show source scenarios. No source-compatible regional queue series.

All measurements 4
Observation485 TWh

Global data centre electricity

The IEA's updated observation for 2025; all data centres, not just AI.

energy/data-centers/2025-electricity
Forecastapprox. 950 TWh

IEA central pathway 2030

Updated scenario, not observed load.

energy/data-centers/2030-updated-case
Interpretation1.96×

Scaling along the central pathway

950 / 485; equivalent to approximately 14.4% compound annual growth.

energy/data-centers/2025-2030-growth
Forecastapprox. 20%

Planned capacity exposed to grid delays

IEA analysis of projects towards 2030; not a project registry.

energy/grid/planned-project-delay-risk
Analysis Data through 2025-12-31
Explore the demand pathway, revisions and grid risks that may constrain compute expansion.

485 TWh observed in 2025; nearly doubling along the IEA's central pathway towards 2030

DEMAND

The IEA's global estimate increased from 415 TWh in 2024 to 485 TWh in 2025.

PATHWAY

The updated IEA pathway reaches 950 TWh in 2030, 1.96× the 2025 level.

REVISION

The central 2030 pathway is almost unchanged between IEA editions: approximately 945 to 950 TWh.

DELIVERY

The IEA estimates that grid risks could delay around 20% of planned global data centre capacity towards 2030.

MEASUREMENT GAP

The module does not yet have a regional registry for power, queue time or connection status.

Full analysis & source chain ↗
VERIFIABLE EVIDENCE / 9 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Calculation

The IEA's updated central pathway reaches 1.96 times the 2025 level in 2030, equivalent to approximately 14.4% compound annual growth.

CALCULATION950 / 485 = 1.9588; (950 / 485)^(1/5) - 1 = 14.4%
  • A mechanical description of the IEA scenario, not an independent forecast.
energy/claim/demand-trajectory
Calculation

The IEA assesses that grid risks could delay approximately one fifth of planned global data centre capacity towards 2030.

  • Model-based risk assessment; not a project registry.
energy/claim/grid-delay-risk
Limitation

The module does not yet have a source-compatible regional series for grid queues, available power or connection status.

  • A measurement gap does not mean regional data does not exist; it has not yet been normalised in this module.
energy/claim/regional-grid-gap
Method

Energy has two compatible global annual observations for total data centre electricity, 2024 and 2025. They are comparable, but two points do not establish a robust long-term trend, and the 2030 points remain source scenarios.

  • The series covers all data centre electricity and is not an AI-only series.
  • Regional grid queues have not yet been measured under the same contract.
energy/claim/evidence-state

Compute × Energy

Data through · 2025-12-31

Compute & energy efficiency

Theoretical compute per nominal MW1,566.5 H100e/MW2025 Q4
Equivalent dense 8-bit capacity3.1 TOPS/W
2025 Q4

Annual electricity model

Sensitivity model
84.3 TWh / year17.4% of 2025 all-data-centre electricity (485 TWh)

Model, not measured consumption: 13.36 GW nominal power × 8,760 h × utilisation × facility overhead. The comparison is not an observed AI electricity share. Useful work per kWh remains unmeasured.

All measurements 4
Interpretation1,566.5 H100e/MW

Theoretical fleet efficiency

3.10 dense INT8 TOPS/W; theoretical proxy.

compute-energy/fleet-efficiency/2025-12-31
Interpretation333.8× / 168.7×

Compute / nominal power

Indexed from 2022-Q1; the difference is the contribution from efficiency gains.

compute-energy/scale/2022-2025
Interpretation117.0 TWh/year

Nameplate envelope

100% utilisation and PUE 1.0; not measured consumption.

compute-energy/nameplate/2025-envelope
Interpretation24.1%

Versus all data centre electricity in 2025

Order-of-magnitude sensitivity against the IEA; not an observed share.

compute-energy/nameplate/2025-data-center-share
Analysis Data through 2025-12-31
See whether capacity growth comes from more power, better theoretical hardware efficiency or both, and test energy assumptions without confusing them with measurements.

1.98× theoretical compute/MW since 2022-Q1; 117.0 TWh/year is only a nameplate sensitivity

SCALE

Compute holdings increased 333.8× while nominal accelerator power increased 168.7×.

EFFICIENCY

Yes, in this peak-performance proxy: 1.98× H100e per nominal MW since 2022-Q1.

ENERGY ENVELOPE

117.0 TWh/year at 100% and PUE 1.0; use the calculator for explicit assumptions.

MEASUREMENT GAP

Not established. This requires workload-specific wall-power and performance measurements.

Full analysis & source chain ↗
VERIFIABLE EVIDENCE / 5 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Calculation

Theoretical fleet efficiency increased from 791.7 to 1,566.5 H100e per nominal MW from 2022-Q1 to 2025-Q4, approximately 1.98 times.

CALCULATIONefficiency = H100e median / nominal accelerator power MW
  • H100e is a theoretical dense 8-bit comparison model.
  • Nominal board power is not actual electricity consumption, and the result is not useful work per joule.
compute-energy/claim/efficiency-trajectory
Calculation

From 2022-Q1 to 2025-Q4, estimated global compute holdings increased approximately 333.8 times, nominal accelerator power 168.7 times and theoretical compute per nominal MW 1.98 times.

CALCULATIONlatest / first for compute, nominal MW and (H100e/MW)
  • The scaling index describes theoretical installed/estimated accelerator capacity, not actual AI work.
compute-energy/claim/scale-versus-efficiency
Calculation

The world series' nominal accelerator power in 2025-Q4 is equivalent to 117.0 TWh per year if nameplate power is sustained continuously at PUE 1.0.

CALCULATION13359.1 MW × 8760 h / 1,000,000 = 117.0257 TWh/year
Epoch AI: Data on AI Chip Owners ↗1 measurement1 source facts
  • This is an upper nameplate sensitivity, not measured energy use.
  • It excludes CPU, memory, networking and cooling, and does not automatically account for PUE above 1.0.
compute-energy/claim/nameplate-energy-envelope
Calculation

The nameplate envelope of 117.0 TWh/year is 24.1 percent of the IEA's estimate of 485 TWh for all data centre electricity in 2025.

CALCULATION117.0257 / 485 × 100 = 24.129%
  • The denominator includes all data centres; the numerator includes only nominal AI accelerator power.
  • The ratio is an order-of-magnitude sensitivity, not an observed share.
compute-energy/claim/data-center-context

AI Research Automation

Data through · 2025-06-27

Research benchmark results

Independent metrics · no combined score

RExBench

12 tasks

Best result with human-written hints

<44 %Research extensions

MLRC-Bench

7 tasks

Share of the gap to top human performance closed

9.3%New ML methods

MLR-Bench

201 tasks / 4 stages

Example of invalid results in one agent setup

Experiment validity risk80 %Invalid experiments

Research workflow coverage

Select a stage
Problem selection

MLR-Bench includes the ideation stage, but within a bounded benchmark framework.

Stage evidence ↗

The sources were published on different dates, but measure different tasks, evaluators and denominators. The module therefore has no defensible longitudinal performance trend yet; the next comparable benchmark release must be added to the same track before drawing a trend.

All measurements 4
Observationbelow 44%

RExBench: research extensions

Best tested result even with human-written hints; 12 tasks.

research-automation/rexbench/best-with-hints
Observation9.3%

MLRC-Bench: gap to top human performance

Best tested agent on seven objectively evaluated competition tasks.

research-automation/mlrc/gap-closed
Observationoften 80%

MLR-Bench: invalid experimental results

Reported for one coding-agent setup; not a universal error rate.

research-automation/mlr/invalid-experiments
Interpretation201 tasks / 4 stages

MLR-Bench: end-to-end coverage

Ideation, proposals, experiments and paper writing; benchmark coverage, not the share of work automated.

research-automation/mlr/workflow-scope
Analysis Data through 2025-06-27
See which R&D steps agents actually demonstrate, and where results still fail objective or experimental checks.

Three separate benchmark perspectives; experimental validity is the clearest bottleneck

IMPLEMENTATION

RExBench v3 reports below 44% even with human-written hints on 12 realistic extensions.

NEW METHODS

The best MLRC-Bench agent closed 9.3% of the gap on seven objectively evaluated competition tasks.

WORKFLOW

MLR-Bench covers 201 tasks and four stages, but reports serious problems with experimental validity.

COMMON SIGNAL

Across different evaluation approaches, implementation and experimental control remain clear weaknesses.

MEASUREMENT GAP

There is no defensible common denominator for an overall automation percentage.

Full analysis & source chain ↗
VERIFIABLE EVIDENCE / 12 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Observation

RExBench v3 reports that the best result, even with human-written hints, is below 44% on realistic research extensions.

Edwards et al.: RExBench paper ↗1 measurement1 source facts
  • Depends on the setup and hints.
research-automation/claim/rexbench-capability
Calculation

Together, the three benchmarks point to implementation and experimental validity as a clear limitation, but they cannot be combined into one automation share.

  • Qualitative triangulation, not a meta-analysis or common score.
research-automation/claim/experimental-bottleneck
Limitation

The module publishes no overall R&D automation percentage because the benchmarks do not measure the same task, agent or success criterion.

  • A measurement gap, not evidence of zero automation.
research-automation/claim/no-aggregate
Limitation

The current benchmarks do not establish a complete, independent validation and reproduction loop after experimentation and interpretation.

  • A coverage gap in the selected benchmarks, not evidence that such validation never takes place.
research-automation/claim/validation-reproduction-gap
Method

The hypothesis stage has partial benchmark coverage: MLR-Bench includes proposals in its workflow, while RExBench defines twelve research extensions as concrete tasks.

  • The sources use different tasks and evaluators and do not establish a common success rate.
research-automation/claim/hypothesis-stage-evidence
Method

Research Automation has three sourced benchmark families, but zero longitudinal points within the same benchmark version. RExBench, MLR-Bench and MLRC-Bench have different tasks, evaluators and denominators and therefore establish no common trend.

  • The benchmark families cannot be combined into one automation share.
  • Maturity describes evidence coverage, not general research autonomy.
research-automation/claim/evidence-state

Agent Task Horizon

Data through · 2026-05-08

Task duration & success rate

Human-expert work time, not how long an agent runs.

30 min2 h8 h32 h
Estimate & source interval16 h reliability boundaryLogarithmic scale

16h · UNRELIABLE ABOVE THIS. Frontier is a reported cohort; named models are separate measurements, not a ranking.

All measurements 6
Observationapprox. 12 h [5–61]

Public frontier, 50% horizon

METR's Feb–Mar 2026 assessment; a wide interval and a suite nearing saturation.

agent-horizon/frontier-2026/public-p50
Observationapprox. 1.5 h [50m–2h40m]

Public frontier, 80% horizon

The same frontier cohort with a stricter reliability requirement.

agent-horizon/frontier-2026/public-p80
Interpretation8×

P50/P80 point estimate

12 hours / 1.5 hours; illustrates sensitivity to the success requirement.

agent-horizon/frontier-2026/reliability-gap
Measurement gap>16 h unreliable

Measurement boundary

METR's explicit boundary for the current suite.

agent-horizon/th11/long-task-boundary
Observation341.735276 min

GPT-5.4, P50

186.581591–768.779526 expert-minutes.

agent-horizon/metr/gpt-5-4/p50
Observation53.877851 min

GPT-5.4, P80

23.957027–108.679232 expert-minutes.

agent-horizon/metr/gpt-5-4/p80
Analysis Data through 2026-05-08
See how much the autonomy horizon shrinks when reliability requirements increase, and when the benchmark runs out of evidence.

Public frontier: approx. 12 h at 50% success, but approx. 1.5 h at 80%

50% SUCCESS

The public frontier is around 12 hours, but the interval is 5–61 hours and extends into the saturation region.

80% SUCCESS

The point estimate falls to approximately 1.5 hours, with an interval of 50 minutes–2 hours 40 minutes.

SENSITIVITY

The rounded P50 and P80 point estimates differ by a factor of eight.

SUITE COVERAGE

31 of 228 tasks take 8+ hours; only five tasks in the entire suite are long tasks with a human baseline.

MEASUREMENT GAP

METR marks horizons above 16 hours as unreliable with the current suite.

Full analysis & source chain ↗
VERIFIABLE EVIDENCE / 10 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Observation

METR TH1.1 reports 320 minutes [170–729] for Claude Opus 4.5 at a modelled 50% success rate.

METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • A specific model, suite and agent setup.
agent-horizon/claim/opus-p50
Observation

METR's February–March 2026 assessment places the public frontier at approximately 12 hours [5–61] at 50% success.

  • A cohort estimate with a wide interval and a suite nearing saturation.
agent-horizon/claim/public-frontier-p50
Observation

The same METR assessment places the public frontier at approximately 1.5 hours [50 minutes–2 hours 40 minutes] at 80% success.

  • A cohort estimate, not a ranking of named models.
agent-horizon/claim/public-frontier-p80
Calculation

The public frontier point estimate is eight times longer at 50% than at 80% success, showing strong sensitivity to the reliability requirement.

CALCULATION12 hours / 1.5 hours = 8
  • The ratio uses rounded point estimates; intervals are wide.
agent-horizon/claim/reliability-gap
Method

TH1.1 has 228 tasks, of which 31 are estimated to take at least eight hours and five of these have a human baseline.

METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • Not a representative distribution of labour-market tasks.
agent-horizon/claim/suite-coverage
Calculation

Only 13.6% of TH1.1 tasks take eight hours or longer, and 2.2% of the entire suite consists of such tasks with a human baseline.

CALCULATION31 / 228 = 13.6%; 5 / 228 = 2.2%
METR: Time Horizon 1.1 ↗1 measurement1 source facts
  • Task counts say nothing about representativeness or quality.
agent-horizon/claim/long-task-coverage
Method

Agent Horizon shows 4 compatible P50/P80 estimates across 2 public model cohorts under the same TH1.1 contract. They are concurrent model and reliability observations, not a time series; suite coverage and the 16-hour boundary limit generalisability.

  • P50 and P80 are reliability levels, not separate points in time.
  • The suite is limited to software, ML and cybersecurity tasks.
agent-horizon/claim/evidence-state
Observation

METR TH1.1 reports 341.735276 expert-minutes [186.581591–768.779526] for GPT-5.4 at 50% success.

  • Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p50
Observation

METR TH1.1 reports 53.877851 expert-minutes [23.957027–108.679232] for GPT-5.4 at 80% success.

  • Model-, suite- and setup-specific estimate.
agent-horizon/claim/gpt-5-4-p80

Advanced Chip Supply

Data through · 2025-12-31
300 mm wafers / month · k = 1,000
2024 · observation850k
2028 · forecast1.4M
ObservedForecast
All capacity points 4
2024 · observation850k
2025 · forecast982k
2026 · forecast1.16M
2028 · forecast1.4M

Dashed segments include source forecasts. Manufacturing capacity is not delivered AI chips. Source unit: thousand 300 mm wafers per month; displayed with k = 1,000 and M = 1,000,000 wafers.

Supply-chain evidence

Different scopes kept separate
Logic / front-endSignal

850k observed in 2024. Compatible 300 mm series; later points are forecasts.

Evidence ↗
HBMData gap

not quantified. No approved compatible public series.

Evidence ↗
Advanced packagingData gap

not quantified. Shares the HBM gap; it cannot be ranked against front-end capacity.

Evidence ↗
Lithography / equipmentData gap

not quantified. No compatible throughput series has been admitted.

Evidence ↗
GeographyForecast

US 10 → 14%. Geographical scenario, not direct AI chip output.

Evidence ↗
Other source units & signals
BROAD ADVANCED CAPACITY · 2.2 M 200mm-eq/month

Different wafer basis; not comparable with the chart.

Evidence ↗
TSMC TOTAL · >17 M 12in-eq/year

Company total; does not isolate AI or leading nodes.

Evidence ↗
US SHARE 2022 → 2032 · 10 → 14 %

Geographical forecast; not direct AI output.

Evidence ↗

HBM + advanced packaging: No approved, compatible public series yet.

All measurements 4
Forecast2.2 M 200 mm-eq./month

Advanced nodes (≤7 nm), 2025

SEMI forecast on a 200 mm-equivalent basis; not actually delivered AI capacity.

chip-supply/semi/advanced-node-2025
Interpretation982k → 1.4 M/month

Advanced 300 mm capacity, 2025 → 2028

Separate SEMI forecast in 300 mm wafers per month; it cannot be added to the 200 mm-equivalent series.

chip-supply/semi/advanced-300mm-path
Forecast<200k → >500k/month

2 nm and below, 2025 → 2028

Leading-node capacity on a 300 mm basis, stated as interval bounds.

chip-supply/semi/two-nm-path
Measurement gapnot quantified

HBM + advanced packaging

No free, method-compatible public series has been approved yet.

chip-supply/hbm-packaging/coverage
Analysis Data through 2025-12-31
See which supply stages have quantified signals, and which critical links still lack compatible data.

Front-end capacity is scaling; HBM and advanced packaging remain measurement blind spots

FRONT-END

Two SEMI publications show strong growth, but use different wafer bases and must not be added together.

LEADING NODES

The source pathway moves from below 200k to above 500k 300 mm wafers per month between 2025 and 2028.

COMPANY

TSMC reports substantial total capacity, but the figure isolates neither AI, leading nodes nor packaging.

GEOGRAPHY

SIA/BCG expects an increased US share, but that share says nothing directly about AI chip output.

MEASUREMENT GAP

No. HBM, advanced packaging and critical equipment still lack approved compatible public series.

Full analysis & source chain ↗
VERIFIABLE EVIDENCE / 10 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Observation

TSMC states that annual capacity at its managed fabs exceeded 17 million 12-inch wafer equivalents in 2025.

TSMC: TSMC 2025 Annual Report ↗1 measurement1 source facts
  • Not advanced-node, AI or packaging capacity.
chip-supply/claim/tsmc-capacity
Limitation

No public, method-compatible series for HBM and advanced packaging capacity has yet been accepted into the module.

TSMC: TSMC 2025 Annual Report ↗1 measurement1 source facts
  • A measurement gap does not mean capacity is absent or that the stage is necessarily the bottleneck.
chip-supply/claim/hbm-gap
Method

The module's 200 mm equivalents, 300 mm wafers, 12-inch annual capacity and geographical shares are kept as separate measures.

  • No hidden area, yield or node conversion is used.
chip-supply/claim/unit-basis
Calculation

Public sources show strong expansion of advanced front-end capacity and leading nodes, while the module cannot yet rank HBM or advanced packaging on the same measurement basis.

  • This is an evidence map, not a complete supply model or bottleneck ranking.
chip-supply/claim/stage-map
Method

Chip Supply has one compatible observed point in the advanced 300 mm series. Three other points are source forecasts and can be read as a SEMI source pathway, not as observed deliveries; HBM, packaging and equipment remain separate measurement gaps.

  • Three forecast points are shown separately from the single observed point.
  • Different wafer, HBM, packaging and equipment measures are explicitly kept separate.
chip-supply/claim/evidence-state