Skip to content
Sidney Scott · The Ashby Institute · July 2026

Via Negativa: The AI Economy by EliminationWhat the constraints permit, 2026 to 2030

Every serious forecast of the AI economy is built the same way. Take a trend, project it forward. Those projections now disagree by two orders of magnitude, because a method that chains assumptions inherits the uncertainty of all of them.

This one is built by subtraction. Enumerate the candidate answers. Compute how long the binding constraint needs to deliver what each candidate requires. Compare against the time remaining. Remove what cannot be delivered. What remains is the forecast.

The constraints are not a checklist. They are derived from a single question that never mentions artificial intelligence: what does it take to obtain one more unit of a limiting input? The answers partition exhaustively into six mechanisms, ordered by how fast each one can supply that unit.

TIER 0Thermodynamicsnever
TIER 1Talent and Absorption10 to 40 years
TIER 2Datanot replenished
TIER 3Matter2 to 7 years
TIER 4Law and Legitimacy3 to 24 months
TIER 5Capitaldays to weeks
Rule weight encodes mutability. Only Tier 0 makes a candidate impossible; the rest make it late, expensive, or contingent.

That ordering does the work, and it produces the first uncomfortable result. Almost all public argument about AI concerns capital. Capital relaxes in weeks. It therefore cannot bind on a multi-year horizon, and an argument about a non-binding constraint cannot change an answer.

The constraint that binds the buildout in 2026 is a transformer. Lead time 128 weeks in 2025, past 160 by 2026, three to five years for the largest units. Not a chip. Not a dollar.

All thirty verdicts by binding constraint, tier and probability
Every verdict, its binding constraint, and its tier. Absorption and Matter bind two thirds of the thirty. Eight are computed against an observable rate; the rest are disciplined judgment, marked as such. The probabilities cluster between 0.60 and 0.80, which is itself a finding: a forecaster whose every answer is probably is barely forecasting.

Contents

Each piece stands on its own. The full arithmetic, all thirty verdicts and the reference list are in the PDF.

I

The Method

Constraints do not eliminate candidates. They reprice them, and only physical law bars outright. The framework is stratified: a deductive floor at Tier 0, an inductive middle across Tiers 1 to 5, and an abductive top across whatever survives. That is why every verdict carries a probability instead of a proof.

II

The Elasticity Gap

Every forecast of AI's contribution to output is a bet on the elasticity between deployed compute and attributable output. Nobody names it, nobody has measured it, and its observed values ran 1.70, then 1.17, then 1.37 across 2023 to 2026. At the low end of that range this paper's own GDP band is eliminated and the sole survivor is the field's most conservative estimate. At the high end even the most bullish survives. The central quantitative dispute in AI economics is a disagreement about one unmeasured number.

III

The Roots and the Economy

AI adds roughly 1.0 to 1.5 percentage points of annual growth, not a majority of GDP: the majority claim fails by about ninefold. Realized gains are running at 0.1 to 0.2 points against a 1.5 point electricity precedent. The binding constraint on all of it is organizational absorption, which relaxes on a twenty to forty year clock.

IV

The Binding Constraints

Compute, data, models, labor, agents, software and the physical world, fifteen questions, one binding constraint each. Models commoditize and the moat moves to everything around them. The data wall is weaker than the field believes. Grid equipment is stronger.

V

The Call and the Record

One dated bet with five adjudication rules: the 80 percent reliability horizon stays under eight hours on 31 December 2027, at P = 0.85. Plus the backtest, including a prediction this framework got wrong at maximum confidence, on the exact constraint tier it had already flagged as vulnerable.

What the engine can decide, and what nothing can

The engine converts a shortfall into a time requirement and a survival likelihood. It decides eight of the thirty questions here. That ratio is reported as a finding, not conceded as a limitation.

It needs two things: a requirement expressible as an expansion factor, and a constraint with an observable rate. Both exist for forward capacity questions and for neither structural questions (does control concentrate, who bears liability) nor current-state ones (are agents reliable today). So this is not a machine for answering everything. It is a test for which questions have physically determinate answers at all. Eight here do. Twenty-two do not, and confident disagreement about those is evidence about the forecasters rather than about the world.

A harder audit, and it cost one of the eight

Asking whether the engine applies is weaker than asking whether it changes anything. Measured against each verdict's own prior, the arithmetic moves eight verdicts by less than 0.12 and is not earning its place. Q8 moves by 0.01. Q14, the data-wall verdict on the tier carrying our recorded miss, moves by 0.07. And Q25 is among them while remaining in the computed set, and both labels are correct. Computed describes its inputs, which are sourced. Decorative describes its output, which does not discriminate. Independent properties.

For those eight the stated probability is substantially a prior wearing arithmetic, and a number that looks computed while being judgment is worse than judgment stated plainly. They are flagged in the repository with the measured shift recorded, so the claim is checkable.

What could be measured, and what could not

This paper set out to measure the migration with a single ratio: value captured by the complements over value captured by intelligence. The intelligence side rebuilt cleanly. The complement side could not be built at all, and reporting that is the more useful result.

Intelligence commoditizes, and faster than we thought. Rebuilt from Epoch AI's replication repository, cloned and re-run rather than quoted: the cheapest model matching GPT-3-class capability fell from $60.00 to $0.07, a fall of 857-fold at 11.8x per year, R² = 0.96. Frontier price fell far less, at 1.34x to 2.25x per year depending on whether reasoning models are allowed to set the frontier. Fixed-capability price falls 4 to 56 times faster than frontier price. That divergence is the commoditization claim, and it needs no index of complements to state.

Frontier and fixed-capability price series and their rates of decline
What the corrected series show. Frontier price falls at 1.34 to 2.25x per year; fixed-capability price falls at 11.8 to 75.3x. The divergence is the commoditization claim, and unlike the withdrawn ratio it needs no index of complements. Epoch AI replication repository, re-run
Withdrawn

An earlier version of this paper claimed frontier price was essentially flat at 1.05x per year, and quoted a Residual Ratio of 990-fold. Both are wrong. The frontier figure was off by up to 114 percent. The ratio had no numerator: none of its four complement components is constructible from disclosed sources. No vendor reports HBM average selling price, no published definition of a data-centre megawatt price exists, TSMC does not break out packaging, and NVIDIA discloses revenue but not units.

Three independent attempts failed for three unrelated reasons, which locates the fault in the construct rather than the instrument. Value capture means rent, rent means price minus cost, and cost is undisclosed at every layer of this stack.

Complement value capture is not measurable from public sources. That is a via negativa finding about the field's own evidence base, arrived at by the method this paper uses on everything else, and it is a more useful contribution than a fourth reconstructed index would have been.

The supply-response test

Every dated claim on this page is tracked in, sixteen entries with six resolved, this paper's own listed first and scored on the same basis as everyone else's. Every citation checked against source is recorded in the verification log, including the sixteen corrections that checking produced.

If capture cannot be measured, the question can be asked another way, and the answer tests the framework's own structure rather than a correlate of it.

A binding constraint attracts no supply response. A cycle does.

Memory is responding. Micron has committed roughly $250bn through 2035 and SK hynix is expanding. High margins are pulling in capital, which is what a cycle looks like, and memory has run this pattern since the 1980s. On this framework's terms that is Tier 5 behaviour: capital converts into capacity.

Power is not. The 2028/29 PJM capacity auction procured 525 MW of new resources, fell 6.8 GW short of the reserve margin target, and cleared with the margin down to 14.7 percent. Four consecutive years of extraordinary prices have induced almost nothing. That is Tier 3 behaviour: capital does not convert inside the horizon.

Same test, opposite answers, both from disclosed sources. This is the first empirical validation of the tier ordering itself. It carries a falsification condition, tracked as signpost S20: if a subsequent auction clears with substantial new entry at or below the cap, the Tier 3 classification of delivered power is wrong.

PJM capacity price signal and supply response: memory vs power
Same test, opposite answers. Memory's margins pulled in $250B of committed expansion. Four years of extraordinary power prices drew 525 MW, 6.8 GW short of target. A binding constraint attracts no supply response; a cycle does. Both answers are from disclosed sources. PJM, company filings
Published because it cuts against us

NVIDIA's gross margin fell to 71.07 percent in FY2026 from 75.0 percent, consistent with rising memory input costs compressing the designer's margin. That is the opposite of what a simple complement-capture story predicts for the accelerator layer.

The intelligence price series and the chip component-cost work both originate with Epoch AI. That is real single-source concentration in the evidence base, and any methodology revision they publish is a revision to this paper.

The Capture Ratio

A second measure asks how much value is captured at all. Against measured US consumer surplus of roughly $172B annually, producers capture about 0.31 of the value created.1 Roughly seventy percent of measured AI value is captured by nobody, accruing to users as surplus on goods priced at zero.

One sourcing note that must travel with this number. Attributable AI revenue, the numerator, rests on three points covering NVIDIA data-centre revenue alone, one of them unverified, with lab run-rates reported rather than filed. The level is indicative; the second decimal is not real.

Value created against value captured, 2025 and 2026
Value created against value captured. Roughly seventy percent of measured AI value accrues to users as surplus on goods priced at or near zero. The hypothesis that capture is collapsing is wrong and withdrawn: K rose from 0.27 to 0.31. The level survives; the trend does not. Brynjolfsson et al. 2026, via the Stanford AI Index 2026
Withdrawn

The hypothesis behind the second unit was that capture is collapsing, competed away faster than any layer can hold it. That is wrong: the ratio rose from 0.27 to 0.31 between 2025 and 2026. The trend claim is withdrawn. Only the level survives.

The evidence, linked

Every load-bearing figure in this paper resolves to a source you can open. Where the evidence is a filing, a price list, a live tracker or something an executive said on a podcast, that is what is linked, rather than a secondary write-up of it.

ClaimEvidence
Interconnection median above 5 yearsLBNL, Queued Up 2026
80% horizon flat at 27 to 32 minMETR, live time-horizon page
29.6 GW AI data-center capacity, Q4 2025Epoch AI database
~300T token public-text stockVillalobos et al., arXiv:2211.04325
Practical CMOS efficiency ceilingHo, Erdil, Besiroglu, arXiv:2312.08595
Power is the constraint, not computeNadella, Bg2 Pod, late 2025
Adoption 88%, agents in single digitsStanford AI Index 2026
Model economics look like commoditiesBOND, Trends: AI
EU AI Act text and deferralRegulation (EU) 2024/1689
US preemption litigation directiveExecutive Order 14365
+78M net jobs, 22% churnWEF Future of Jobs 2025
5% of pilots P&L-positiveMIT NANDA, GenAI Divide
Verification should cost one click. A paper that asks you to accept a lead time or a price without showing you where it came from is asking for trust it has not earned.
  1. Consumer surplus from Brynjolfsson et al. (2026), longitudinal willingness-to-accept estimates, reported in the Stanford AI Index 2026: $112B rising to $172B annually, median value per user tripling from $3.40 to $11.40. Captured value is attributable AI revenue. Because the surplus figure covers US consumers only and excludes enterprise surplus entirely, the true capture ratio is lower than stated.
Piece I

The Method

Constraints do not eliminate candidates. They reprice them, and only physical law bars outright. The framework is stratified: a deductive floor at Tier 0, an inductive middle across Tiers 1 to 5, and an abductive top across whatever survives. That is why every verdict carries a probability instead of a proof.

What the maxim licenses

When you have eliminated the impossible, whatever remains, however improbable, must be the truth.

The maxim is deductive, and it is valid only where the candidate set is exhaustive and each elimination is a genuine impossibility. Neither condition holds for most questions about the AI economy. Candidate sets over markets are not provably exhaustive, and almost nothing in economics is impossible.

A framework claiming uniform deductive force will fail the way every constraint-based forecast has failed, by treating an improbability as an impossibility. Malthus did it. The Ehrlich side of the Simon-Ehrlich wager did it. So this framework is stratified by inferential type.

LayerOperationApplies atVerdict
Deductive floorElimination of the impossibleTier 0 onlyBarred
Inductive middleRepricing against observed ratesTiers 1 to 5Repriced
Abductive topInference to the best remainderAcross survivorsIsolated

The probabilities were always the tell. A framework that declares candidates eliminated and then attaches survivor probabilities of 0.55 to 0.78 is describing itself wrongly: a survivor at 0.60 leaves forty percent of the mass with candidates it called eliminated. That is incoherent under a deductive reading and coherent under a stratified one.

Three inferential layers: deductive floor, inductive middle, abductive top
Only the floor licenses deduction. Constraints do not eliminate candidates; they reprice them. A survivor at 0.60 leaves forty percent of the mass with candidates the framework called eliminated.

Deriving the six

The constraint set is open to a circularity charge if it is chosen with the conclusion in view. So it is derived from a question that never mentions AI: what does it take to obtain one more unit of a limiting input?

Source of one more unitConstraintAccessible in
Nothing. Physical law forbids itThermodynamicsnever
Generational time: humans raised, trained, organizedTalent and Absorption10 to 40 years
Nothing, but a finite accumulated stock is drawn downDatanot replenished
Industrial production, given capital and lead timeMatter2 to 7 years
A collective decision by people with authorityLaw and Legitimacy3 to 24 months
A priceCapitaldays to weeks
The partition is over replenishment mechanisms, not subject matter, and it is complete. An input is either forbidden, biological, fossil, manufactured, permitted, or purchased. The tier ordering falls out of the derivation rather than being asserted beside it.

Three placements follow from the derivation rather than from preference. Data is a fossil resource, an accumulated stock with no meaningful replenishment, which places it above Matter despite being informational. Absorption is Tier 1 because it is biological: the capacity of firms to reorganize work turns over with people and management generations. And energy is not a primitive; it is an output of Matter and Law, floored by Thermodynamics.

That last one matters. Treating energy as one constraint conflated a physical invariant with two orders of magnitude of headroom, an industrial stock with none, and a permitting regime that moves in months. On this horizon delivered electricity binds through transformer lead times and interconnection queues, never through physics.

The engine

Every test reduces to one comparison: time required against time available. For candidate a and constraint c, the slack ratio is the quantity permitted over the quantity required, σ = Permitted / Required, and the shortfall is g = 1/σ. The time the constraint needs to deliver that expansion at its observed rate r is

Treq  =  ln(g) / ln(1 + r)

and survival is a logistic in the difference between time available and time required, with S = 0 whenever the constraint is Tier 0. That zero is the deductive floor, and it is the only place a zero appears.1

Required time as a function of shortfall, and the survival logistic
The engine. Left: how long a constraint needs to close a shortfall at its observed expansion rate, against the horizon remaining. Right: survival as a logistic in the difference. The only zero in the system sits at Tier 0.

The parameters are observable rather than chosen, which is the substantive improvement over a penalty-coefficient approach. Matter is decomposed by sub-constraint because its expansion rates differ by an order of magnitude: advanced packaging has expanded at 60 to 110 percent annually while grid equipment has expanded at 10 to 20. That ratio alone explains why the binding constraint migrated from packaging in 2024 to grid equipment in 2025 and 2026, with no additional assumption required.

Observed annual expansion rates by constraint and tier
Why the binding constraint migrates. Sub-constraint expansion rates differ by an order of magnitude, so advanced packaging absorbs shortfalls that grid equipment cannot. The ratio between those two rates is the entire content of the migration argument.
Where this does not apply

The engine needs a candidate whose requirement expresses as an expansion factor and a constraint with an observable rate. Both exist for forward capacity questions. Neither exists for structural questions (does control concentrate, who bears liability) nor current-state ones (are cash flows sufficient today). Eight of thirty verdicts are computed. The rest are disciplined judgment, and each one says which it is.

Eight of thirty verdicts computed, fourteen structural, eight current-state
Where the engine applies. It needs an expansion factor and an observable rate. Both exist for forward capacity questions and for neither of the other classes.

Where time enters

Time is not a constraint. It has no independent permitted quantity, and one cannot compute time permitted over time required without first naming the process being timed. But three genuine time-dimension mechanisms exist, and none is captured by a single slack ratio.

Allocation is not stock. Every constraint has two clocks that differ by one to two orders of magnitude. A firm reserving more than half of a foundry's packaging expansion shortens no lag; it establishes queue position. An allocation mechanism never relaxes a constraint at the system level, only at the actor level, so system questions and firm questions are tested against different denominators.

Synchronization. Electricity is not storable, so a site needs power at the coincident peak rather than on average. The relaxation direction is the interesting one: a candidate that accepts interruption converts a synchronization requirement into an energy requirement, and curtailable interconnection clears in months where firm service clears in years. Flexible load is the single largest available relaxation lever on the buildout.

Serial path. A candidate can fail with every individual slack ratio above one, because the chain of conversions it requires is longer than the horizon. A greenfield site with no interconnection agreement needs six to seven years of serialized conversions against 4.4 remaining to 2030. That is a timing verdict, not a possibility verdict, and it must say so. Conflating the two is how constraint forecasting earns its reputation for crying impossible when it meant late.

Conversion lags from capital to each constraint against the horizon
Money becomes megawatts, but only after the lag. Everything to the right of the line cannot be bought into existence before 2030 at any price.

The parameters, in full

These are the numbers every verdict runs on. Change one and the verdicts change, which is the sense in which this forecast is auditable rather than merely sourced.

TierConstraintrcGoverning datum, mid-2026
0Thermodynamics0 exactlyLandauer bound ~3e-21 J/bit at 300K. Max CMOS ~4.7e15 FP4/J, roughly 200x current accelerators
1Talent0.05 to 0.10Inbound AI researchers to the US down 89% since 2017, 80% in the last year alone
1Absorption0.02 to 0.05Electric motors under 5% of factory drive in 1900, gains in the 1920s. ~5% of AI pilots P&L-positive
2Data0.30 to 0.60~300T tokens effective public text, median exhaustion 2028. Revised upward by an order of magnitude after the backtest miss
3aMatter: packaging0.60 to 1.10CoWoS from ~13k wafers/month end-2023 toward a 120 to 130k target end-2026
3bMatter: logic fabs0.15 to 0.25TSMC Arizona announced 2020, high-volume production late 2024 to 2025
3cMatter: grid equipment0.10 to 0.20Transformers 128 weeks in 2025, past 160 by 2026. ~80% imported. Switchgear sold out through 2028
3dMatter: interconnection0.15 to 0.25LBNL median request to operation above 5 years for 2025 projects; ~2,061 GW in queues
4Law and LegitimacystepEU high-risk duties deferred 16 months by the Digital Omnibus, Council approval 29 June 2026
5Capital0.362026 hyperscaler guidance ~$690B to $725B. The $1.5T is a funding gap, not a debt requirement
The slowest rate on a candidate's path is the one that decides it. Grid equipment expands at roughly a tenth the rate of advanced packaging, which is the whole content of the migration argument and requires no further assumption.

Find the binding constraint

The operational instruction is not "check all six." It is three steps, and it is portable to any claim you encounter.

Screen. Order-of-magnitude every constraint. Most resolve in one line. Compute. Take the two or three lowest and do them properly, with sourced numerators and denominators, as intervals rather than points. Attribute. Name the binding constraint, state its tier, state the gap, and state the discontinuity that would move it.

A verdict that does not name a single binding constraint has not been resolved. The reason this matters beyond bookkeeping is a result from linear programming: a non-binding constraint has a shadow price of zero. An argument about a non-binding constraint cannot change the answer. Most public forecasting of the AI economy argues about capital, which on this horizon almost never binds.

  1. Full form, parameters, conversion lags and the treatment of uncomputable cells are in the PDF, Section 1. An uncomputable cell does not receive a free pass: it takes a base-rate-bounded interval, and the verdict reports the posterior at both ends. Simulation showed the earlier free-pass treatment was carrying two thirds of an unmeasurable candidate's posterior mass.
Piece II

The Elasticity GapThe parameter the whole field is betting on

Every forecast of AI's contribution to output is a bet on the elasticity between deployed compute and attributable output. Nobody names it, nobody has measured it, and its observed values ran 1.70, then 1.17, then 1.37 across 2023 to 2026. At the low end of that range this paper's own GDP band is eliminated and the sole survivor is the field's most conservative estimate. At the high end even the most bullish survives.

The parameter nobody names

Every forecast of AI's contribution to output implicitly assumes a relationship between deployed compute and attributable output. Write it down:

O  =  k · Cα

where O is AI-attributable output, C is deployed compute, and α is the elasticity of the first with respect to the second. At α = 1 output scales proportionally with compute. Below one, each additional gigawatt yields less than the last. Above one, deployed compute yields increasing returns, as it would if a model trained once serves many users or if value accrues through channels consuming little marginal inference.

No published forecast of the AI economy states its α. None reports it, none defends it, and it has never been measured. Yet every such forecast is a bet on its value, because the forecast asserts an output figure, the constraints permit a quantity of compute, and only α connects them.

Published forecasts against the Tier 3 constraint as a function of alpha
Every published forecast, run through the constraint engine. At α = 1.0 this paper's own band is eliminated and only Acemoglu survives; at α = 1.4 both survive. The shaded band is the range actually observed. Drag the slider below, or reproduce from the repository
Drag it yourself

Every forecast below is a bet on this one number. Nobody has measured it. Move the slider and watch which survive.

α = 1.37observed, 2025 to 2026
Observed values: 1.70 (2023 to 2024) · 1.17 (2024 to 2025) · 1.37 (2025 to 2026). Anchors: $110B attributable output on 29.6 GW in 2026; 200 GW permitted to 2030; 40% annual efficiency gain.

Every forecast, run through the engine

Anchoring on 2026 at roughly $110B of attributable output on 29.6 GW of AI data-center capacity, taking permitted incremental capacity to 2030 as roughly 200 GW, and applying the standard efficiency correction of about 40 percent per year, the slack ratio for each published forecast is a function of α alone.1

Forecastα=0.50.71.01.21.41.7
PwC 2017, +$15.7T by 20300.000.020.180.420.751.40
Goldman 2023, +7% global GDP0.000.040.260.570.981.75
McKinsey 2023, $4.4T annual0.020.130.651.201.862.96
This paper, upper $10.5T0.000.040.270.581.001.78
This paper, lower $7.0T0.010.070.410.821.342.26
Acemoglu 2024, 0.66% TFP over 10 yr0.110.521.692.663.685.20
Cells give σ, permitted over required. A forecast survives at σ ≥ 1, shown in teal; struck values are eliminated by the Tier 3 constraint.

The observed values of α, computed from the attributable-output and capacity series, are unstable and sit in the increasing-returns region:

PeriodOutputCapacityImplied α
2023 to 20245.00x2.57x1.70
2024 to 20252.40x2.11x1.17
2025 to 20261.83x1.56x1.37

What the table says

At α = 1.0, this paper's own GDP band is eliminated. Both bounds fail, at σ = 0.27 and σ = 0.41, alongside Goldman and PwC. The single survivor is Acemoglu, the most conservative published estimate in the field.

At α = 1.4, the midpoint of the observed range, this paper survives and so does Goldman. At α = 1.7, the value observed across 2023 to 2024, even PwC survives.

So the central quantitative dispute in this field is a disagreement about α that no participant has named. PwC's $15.7T is a bet that α ≥ 1.7. Goldman's 7 percent is a bet on roughly 1.4. Acemoglu's 0.66 percent is a bet that α ≤ 1.0. The verdicts in this paper are a bet on roughly 1.3 to 1.4. None of them has stated the wager it is making.

Implied elasticity 2023 to 2026: 1.70, 1.17, 1.37
The elasticity is not a stable parameter. It ran 1.70, then 1.17, then 1.37. That range is wide enough to reverse every forecast in the table above, and no participant in the debate reports it.

The observed range spans values that reverse every forecast in the table. That is not a narrow uncertainty around a central estimate. It is the difference between a technology that adds one percent of output and one that adds ten.

What it costs this paper

Conceded

The GDP verdict cannot be resolved at the confidence a single-point analysis suggests. Its probability falls from 0.62 to roughly 0.55 and is marked α-dependent and provisional. Under α = 1.0 it is not resolved at all.

The temptation is to select the α that preserves the band and proceed. That would be the characteristic failure this framework was built to avoid, committed on the framework's own numbers: holding a parameter fixed at a convenient value while the evidence says it is unstable.

The exposure is concentrated rather than general. Verdicts about output magnitude are bets on α. Verdicts about physical delivery, lead times, market structure, liability and reliability are not, because they never require the compute-to-output conversion. Each is marked where it applies.

What would settle it

A panel of AI-attributable output against deployed compute at firm or sector level, with output measured as realized revenue plus estimated surplus rather than as capital expenditure. That series does not exist. Building it is the single most valuable unbuilt dataset in the economics of AI, and it would settle a dispute currently conducted entirely through unstated assumptions.

The point of this piece

This paper does not add the forty-seventh forecast. It eliminates the claim that the question is currently answerable, and names what would answer it. That the result costs the author a headline number is the strongest available evidence that the method is doing work rather than decorating a conclusion.

  1. Anchors: attributable revenue and capacity series in the PDF, Section 2; 29.6 GW of AI data-center power capacity at Q4 2025 from Epoch AI via the Stanford AI Index 2026; permitted incremental capacity to 2030 from Bain's 6th Global Technology Report; efficiency gains from Epoch AI.
Piece III

The Roots and the Economy

The majority-of-GDP claim is dead: it needs $77T of new output by 2030 and the most bullish credible estimate permits $16T. What survives is 1.0 to 1.5 points of annual growth. Realized gains today are 0.1 to 0.2 points against a 1.5 point electricity precedent. Absorption binds, and it relaxes on a twenty to forty year clock.

The five roots

Beneath the hundred questions lie five root uncertainties. Every branch question inherits its verdict from one or more of them.

R1. Will AI create more value than it destroys?

Aggregate value destruction is eliminated at the sign level: every prior general-purpose technology delivered a positive contribution once diffused. What survives is a J-curve. Destruction front-loads, creation lags, and the capital gap dates the lag rather than the destination: $725B of 2026 capex against $110B of attributable revenue.

binding: Absorption, Tier 1 · P ≈ 0.73 · tails: distributional rejection 0.15, financing cascade 0.07

Hyperscaler capex against attributable AI revenue, 2023 to 2026
The ratio dates the lag, not the destination. Hyperscaler capex ran at 30x attributable AI revenue in 2023, converging to 6.6x by 2026. It is a funding gap and not a debt requirement: reading the widely cited $1.5T as debt overstates the capital constraint roughly sevenfold, on the one constraint least likely to bind. Company filings, SEC

R2. How fast do organizations become AI-native?

Overnight transformation requires broad P&L-positive deployment now. Five percent of integrated pilots are P&L-positive. That is a twentyfold gap on the success base, and no mechanism closes it in three years. A decade-long uneven grind survives. This is also where Absorption was discovered: the bottleneck is organizational, and a constraint set without it would find that and have nowhere to put it.

binding: Absorption, Tier 1 · P ≈ 0.70

R3. Who controls compute, energy, data and models?

Monopoly is barred on structure: no actor owns accelerators, foundry, cloud and models. Decentralization is barred on delivered power. What survives is a US-led layered oligopoly beside a walled Chinese stack. Concentration is real and split across layers, and nobody owns the floor, which is one foundry in Taiwan and the public grid.

binding: Matter, Tier 3 · P ≈ 0.75

Market shares by layer: packaging, accelerators, foundry, cloud
Concentration is real and split across layers. Each layer is concentrated and owned by a different firm, which bars the monopoly candidate on structure. Nobody owns the floor.

R4. Can institutions adapt without stifling innovation?

Fragmented, lagging, sectoral governance. The operative risk is divergence and arbitrage rather than stringency.

binding: Law and Legitimacy, Tier 4 · P ≈ 0.78

R5. What are the real limits of intelligence?

Scale-to-AGI-soon is eliminated, but not for the reason the field gives. The pre-training data wall is no longer load-bearing here. Another hundredfold of training compute requires fab and grid expansion that Tier 3 cannot deliver by 2030. Capability keeps rising through test-time and post-training compute. The lever moved; it did not break.

binding: Matter, Tier 3 · P ≈ 0.65 · jagged plateau as near-term texture, 0.30

A demonstration of Tier 4

The EU AI Act's high-risk obligations, treated as fixed at 2 August 2026 for over a year, were deferred to 2 December 2027 and 2 August 2028 by the Digital Omnibus, with final Council approval on 29 June 2026. A flagship compliance deadline moved sixteen months in seven months of legislative process. No Tier 3 constraint can move like that and no Tier 0 constraint can move at all.

How much GDP

The majority-of-GDP claim requires AI-attributable output above half of a $154T 2030 economy. That is $77T of new output in four years. The most bullish credible estimate in the field, PwC's $15.7T cumulative, permits $16T. A ninefold gap, and no constraint in the set closes it in four years.

Majority-of-GDP claim against the feasible band
The majority claim against the feasible band. The survivor sits inside the electricity and information-technology precedents rather than above them.

What survives is 1.0 to 1.5 points per year, $7 to $10.5T cumulative by 2030. That is inside the electricity and information-technology precedents, not above them. The bulls are not wrong about the technology. They are wrong about the denominator.

Two qualifications, both of which narrow the claim. It is α-dependent and provisional. And it forecasts measured GDP: value delivered as consumer surplus on goods priced at zero sits outside that quantity by construction, and US consumer surplus alone is running at roughly $172B annually and growing 54 percent. A reader concluding that AI creates little value from a modest GDP verdict has misread it. The verdict is about capture and measurement, not welfare.

Why it arrives slowly

An instant productivity surge is not a live option. It requires 69.9 years of absorption expansion against 4.4 available. Internet-speed diffusion is not one either: the binding constraint is deployment rather than distribution, and the precedent is two to four decades. Electric motors were under five percent of factory mechanical drive in 1900. The productivity gains landed in the 1920s.

The distinction that matters, and that headline adoption figures obscure: consumer uptake and organizational absorption run on different clocks. Generative AI reached 53 percent population adoption in three years, faster than the PC or the internet. Over the same period AI agent deployment stayed in single digits across nearly all business functions. Both facts are from the same report.1

The strongest contrary datum

This could falsify the verdict

US labor productivity growth reached 2.7 percent in 2025, nearly double the 1.4 percent average of the preceding decade, which Brynjolfsson reads as the early phase of a J-curve. If that is AI-attributable and sustained, the electricity-slow verdict is wrong.

Three considerations bear on it and none is decisive: aggregate productivity growth is not AI-attributable by default; the J-curve reading is this paper's own thesis, since its early phase is absorption cost preceding measured gain; and one year against a decade average is within the historical variance of the series.

Falsification threshold, stated in advance: US productivity growth holding above 2.5 percent for three consecutive years with a credible AI attribution falsifies this verdict. Tracked as signpost S18.

The rest of the domain

Q4. How much new wealth, and to whom?

Trillions, concentrated, accruing to the complements rather than to intelligence itself. Carry the full Acemoglu-to-McKinsey order-of-magnitude range rather than a point.

binding: Capital, Tier 5 · P ≈ 0.80

Q5. Which industries gain most?

Digital, measurable, verification-cheap sectors first. Micro evidence tracks the gradient precisely: +14 to 15 percent for support agents, +26.08 percent (SE 10.3) for developers using Copilot, and -19 percent for experienced open-source developers, who became slower while believing they had been sped up.

binding: Data, Tier 2 · P ≈ 0.80

Q6. Which industries become obsolete?

None this horizon. The routine middle compresses: commodity outsourcing, generic content, tier-1 support, basic data processing.

binding: Law, Tier 4, with Absorption · P ≈ 0.80

Q7. Inequality?

Rises through the transition absent deliberate redistribution. Cheap intelligence does not democratize the scarce complements.

binding: Capital, Tier 5, with Law · P ≈ 0.75

Q8. Jobs?

Roughly flat-to-positive by count, +78M net by 2030, with 22 percent churn. The pain is in the transition: the destroyed jobs are not the created ones. Employment for software developers aged 22 to 25 has already fallen nearly 20 percent from 2024.

binding: Capital, Tier 5 · P ≈ 0.70

Q9. How long does the transition take?

A decade and more, clocked by deployment friction and grid interconnection rather than model capability. Only an agent-reliability inflection resets it.

binding: Matter 3d and Absorption · P ≈ 0.72 · computed: 46.5 years required against 2.0 available

Q10. Growth or reallocation?

Reallocation and margin first, real growth later, with a live investment-incineration tail if GPU infrastructure pricing competes to marginal cost.

binding: Capital, Tier 5 · P ≈ 0.65 · incineration tail 0.25

  1. Adoption and agent deployment both from the Stanford AI Index 2026, which also supplies the micro-study gradient in Q5 and the entry-level employment figure in Q8. Productivity figures and the Brynjolfsson J-curve reading are from the same source.
Piece IV

The Binding ConstraintsCompute, data, models, labor, agents, software, the physical world

Fifteen questions, one binding constraint each. Models commoditize and the moat moves to everything around them. The data wall is weaker than the field believes. Grid equipment is stronger. The chokepoint on the buildout is a transformer with a 128-week lead time, not a chip and not a dollar.

Compute and its financing

The chokepoint, dated

US large-power transformer lead times average 128 weeks and reach three to five years for the largest units. Medium-voltage switchgear is effectively sold out through 2028. The gas-turbine order book reached 100 GW with roughly 10 GW of delivery slots remaining across 2029 and 2030 combined. The LBNL median from interconnection request to commercial operation exceeded five years for projects built in 2025.

Of 12 to 16 GW of US data-center capacity announced for 2026, roughly 5 GW is under construction. Closing that gap requires 5.2 years of expansion at the observed rate against 1.4 years available.

Announced US 2026 data-center capacity against capacity under construction
The binding constraint, dated to the week. Of announced US capacity for 2026 delivery, roughly a third is under construction. Closing the gap needs 5.2 years of expansion at the observed rate against 1.4 available. LBNL, Wood Mackenzie, GE Vernova

Q11. Who controls compute?

A layered US-led oligopoly with a live custom-silicon erosion tail. One foundry fabricates nearly every leading AI chip, which is the concentration that matters.

binding: Matter, Tier 3 · P ≈ 0.75

Q12. How much compute will AI need?

Unlimited is eliminated at Tier 3, and the mechanism is dated to the week. The ceiling is grid equipment and interconnection. It is not appetite, not capital, and emphatically not physics: accelerators sit two orders of magnitude below the practical CMOS ceiling. Anyone claiming AI faces a near-term physics wall is making a Tier 3 claim and mislabelling it.

binding: Matter 3c, Tier 3 · P ≈ 0.75 · Treq 5.2 yr vs 4.4 available

Q13. Who finances the next trillion?

Debt and private credit, not cash flow. But the widely cited $1.5T is a funding gap, not a debt requirement, and reading it otherwise overstates the capital constraint roughly sevenfold. Stress is idiosyncratic before it is systemic.

eliminated by observation, not constraint · P ≈ 0.78 · cascade tail 0.10

Data

Q14. Does AI run out of data?

Public text exhausts around 2028 on the median and scaling migrates to verified synthetic, multimodal and private sources. The scarce complement moves from raw data to verification apparatus rather than disappearing.

binding: Data, Tier 2 · P ≈ 0.60, reduced · confidence bounded by the backtest miss

This is the verdict the paper is least confident about, and the reason is in Piece V: the framework predicted the data wall would bind by 2026, at maximum confidence, and was wrong. Tier 2 is where substitution defeats stock-based reasoning.

Models and scale

Q15. Do models commoditize?

One model winning is eliminated: no frontier lead has held for more than months. Zero margin everywhere is eliminated: the newest model always commands a premium. What survives is that the floor commoditizes fast and the frontier holds a brief, eroding lead. The model is a commodity and the moat is everything around it. BOND reaches the same premise and calls the consequence a riddle. It is not a riddle. The floor falls 11.8x to 75.3x per year while the frontier falls 1.34x to 2.25x, and that divergence is the whole answer.

Frontier and fixed-capability price series and their annual rates of decline
The two prices, measured. Frontier price falls 1.34x to 2.25x per year depending on whether reasoning models set the frontier. The price of a fixed capability falls 11.8x to 75.3x. Epoch AI replication repository, re-run

binding: Capital, Tier 5 · P ≈ 0.75

Q16. Does scale still work?

Yes, but the lever moved, and it moved because of fabrication and power rather than tokens. The question flips from how big to pre-train to how much thinking to buy at inference.

binding: Matter 3b, Tier 3 · P ≈ 0.68

Two denominators, and they are not interchangeable

Fixed-capability inference price falls roughly tenfold per year; Stanford HAI measures a 99.7 percent fall in cost per token between November 2022 and December 2024, about 15.9x annually. The frontier price, the best model at launch, falls at roughly 2.9x. The Residual Ratio uses the frontier denominator. Under the fixed-capability denominator it is larger by orders of magnitude. Both are legitimate; using them interchangeably is not.

Labor and value

Q17. How much knowledge work is automatable?

Task-level, not job-level. Exposure is broad but cost-effective automation is partial and rising. Note the attribution: the 23 percent figure is a cost line, which is Tier 5 and relaxes in weeks, so this elimination is weak.

binding: Capital, Tier 5 · P ≈ 0.75

Q18. What gains value when intelligence is cheap?

The complements intelligence cannot make abundant: energy, verified data, fabrication capacity, trust, distribution, and human accountability for consequential decisions. The thesis in its purest form.

no binding constraint on the survivor · P ≈ 0.80

Agents

Q19. Can agents be trusted to act reliably?

Already-reliable is eliminated by measurement: production single-task success clusters near 56 percent and roughly 88 percent of pilots never reach production. Never is eliminated by the trend. Agents are trustworthy inside bounded, verifiable, short-horizon tasks and nowhere else. Trust scales with verifiability, not with raw capability.

empirical, not constraint · P ≈ 0.70 · sub-threshold plateau 0.25

Q20. Can agents run whole workflows unsupervised?

No, and the limit is arithmetic rather than sentiment. Success compounds multiplicatively. At 95 percent per step, a twenty-step workflow succeeds 36 percent of the time. Long unsupervised chains need 99 percent-plus per step. Production is at 56.

Whole-task success against workflow length at several per-step rates
Why long unsupervised chains fail. The limit is arithmetic rather than sentiment, which is why this verdict is labelled empirical and not constraint-based.

empirical, not constraint · P ≈ 0.72

Q21. Who is liable for an agent's decisions?

The deploying principal, by default. The residual responsibility gap for genuinely autonomous acts makes accountable human oversight a scarce, valuable complement: the Residual Ratio appearing in law.

binding: Law, Tier 4 · P ≈ 0.75

Software and the physical world

Q22. Does SaaS survive?

It survives but reprices, from seats to outcomes and data, with margins bifurcating.

binding: Capital, Tier 5 · P ≈ 0.70

Q23. When are humanoid robots viable?

In structured, high-wage, multi-shift niches first. The last-ten-percent dexterity gap that keeps general deployment out is task horizon and scope, not simulation-to-reality. On BEHAVIOR-1K, a benchmark of long-horizon mobile manipulation in simulated homes, the top team reached 12.4 percent full-task success; on the short-horizon 18-task RLBench subset, results reach the high eighties. Both are simulation. Performance collapses with horizon and scope, and a household is exactly that.

binding: Matter, Tier 3 · P ≈ 0.60, computed · mass deployment retains 0.23

Q24. When do autonomous vehicles go mainstream?

City by city, gated by per-jurisdiction permission and validation. Mainstream-everywhere is a decade-plus away.

binding: Law, Tier 4 · P ≈ 0.72

Q25. Does intelligence become an abundant utility?

Free and infinite is eliminated: unbounded supply needs eleven years of Tier 3 expansion against 4.4 available. Stays-scarce is eliminated by the price collapse. Both survivors say the same thing. Tokens become utility-cheap and the scarce utility is the layer beneath: power, compute, fabrication. The Residual Ratio reaches its maximum here.

binding: Matter, Tier 3 · P ≈ 0.78

Piece V

The Call and the Record

One dated bet with five adjudication rules: the 80 percent reliability horizon stays under eight hours on 31 December 2027, at P = 0.85. Plus the backtest, including a prediction this framework got wrong at maximum confidence, on the exact constraint tier it had already flagged as vulnerable.

The Call

The headline call

Reliability, not capability, remains the binding constraint on AI agents through 2027. On 31 December 2027, the best generally available model's METR 80 percent-reliability task-completion time horizon is under eight hours.

P = 0.85 · adjudicated 31 March 2028 · signpost S13

The widely cited METR figure is the 50 percent horizon, and it has risen several-fold since late 2025. The 80 percent horizon, the level at which work can actually be delegated, has been approximately flat at 27 to 32 minutes across two frontier release cycles.1

Falsification therefore requires roughly a sixteen-fold rise, four doublings, in under two years, in a series that has not moved. The mechanism is the argument rather than the extrapolation: capability gains are arriving as intercept rather than slope. Models are becoming better at succeeding sometimes on long tasks without becoming better at succeeding reliably.

50 percent horizon rising while the 80 percent horizon stays flat
The divergence the Call is staked on. The cited 50 percent horizon has risen several-fold. The 80 percent horizon, the level at which work can be delegated, has been flat at 27 to 32 minutes across two frontier release cycles. METR

Adjudication rules

A bet that cannot be settled is not a bet.

  1. Series. METR Time Horizon 1.1 or its successor, as published at metr.org/time-horizons.
  2. Methodology change. If METR replaces the series, the Call adjudicates on the successor. If the successor reports no 80 percent horizon, the Call resolves void rather than in either party's favour.
  3. Generally available. Any member of the public can obtain access by payment, without individual negotiation, waitlist approval or preview enrollment.
  4. Unmeasured models. Adjudicated on the highest published 80 percent horizon among generally available models. If the strongest has no published horizon, the Call resolves on those that do and the omission is recorded.
  5. Adjudication date. 31 March 2028, to allow for publication lag.

Two hazards acknowledged: METR has noted that measurements above 16 hours are unreliable with the current task suite, and its coverage is explicitly not comprehensive. Both make the Call harder to settle rather than easier to win.

The backtest

A framework that has never been scored is falsifiable in principle and not in practice. Five questions were pre-registered with parameters fixed at their historical as-of dates, before outcomes were consulted. One was chosen because the framework predicts its own failure there.

QuestionAs-ofP(yes)ActualBrier
CoWoS packaging meets 2024 demandend-202336.4%NO · hit0.132
Capital is binding on 2025 buildoutend-20230.0%NO · hit0.000
Data wall binds scaling by 2026end-2022100.0%NO · MISS1.000
US capacity meets 2025 announcementsend-20238.7%NO · hit0.008
Grid equipment binding on 2026 buildoutend-202497.8%YES · hit0.001
Four of five correct, mean Brier 0.228 against 0.25 for a coin flip. The aggregate is dragged there entirely by one catastrophic miss; removing it gives 0.035.

The miss

Wrong at maximum confidence

The framework predicted, at 100 percent confidence, that the public-text data wall would bind frontier scaling by 2026. It did not. Epoch's own median exhaustion estimate moved from 2024 to 2028, and multimodal corpora and verified synthetic generation substituted for the stock the model had treated as fixed.

This is the classical failure of resource-limit forecasting, reproduced inside our own model, on the tier we had already flagged as exposed to it. Malthus, the Ehrlich side of the Simon-Ehrlich wager, and The Limits to Growth against the Nordhaus critique all failed identically: the analyst held Permitted fixed while the world held it variable. We named the mechanism and then failed to parameterize it.

Backtest of five pre-registered questions, four hits and one miss
The record, including the miss. The vertical bar is the outcome. The single miss came at maximum confidence on Tier 2, the tier this framework had already named as most exposed to substitution.

Three corrections followed. The Tier 2 expansion rate was revised by an order of magnitude, from 0.02 to 0.05 up to 0.30 to 0.60, because the parameter must be the rate at which the effective stock expands including substitution. Three verdicts moved. And the Brier score is now reported by tier rather than in aggregate: the framework is excellent on Tiers 3 and 5, where constraints are industrial and rates are observable, and it was worthless on Tier 2.

What the backtest does not establish: five questions is not a validation set, and four of five are physical-infrastructure questions, the domain where the method should perform best. It is untested on institutional, labor and behavioural questions, which is most of the economic domain, and is probably weaker there.

The objection that is granted

Six objections are answered in full in the paper. One is granted outright.

Granted

Every historical analogue this paper invokes has the same structure: the final binding constraint was organizational, not physical. In electrification, generation and distribution were solved decades before the productivity gains arrived. In the 1990s telecom buildout the constraint moved from capital to demand, with roughly 5 percent of installed fiber lit by 2001.

If the same holds here, the residual accrues to whoever solves organizational adoption rather than to owners of physical complements, and the investment conclusion inverts. The thesis is therefore narrowed: it holds on the 2026 to 2030 horizon, where physical constraints dominate and absorption has not had time to bind. Beyond that horizon this paper does not make the claim.

The register, in full

Nineteen tracked signposts: 8 constraint, 5 state, 5 coupling, 1 meta
What the register measures. Five of the original thirteen measure the world moving inside the walls rather than a wall moving. A register that does not distinguish the two invites over-reading of every trigger.

No participant in this field maintains a track record. Forecasts are published, cited, and quietly superseded. This is the scoreboard, and this paper's own entries are listed first and scored on the same basis as everyone else's.

Resolved: the pre-registered backtest

QuestionAs ofP(yes)OutcomeBrier
CoWoS meets 2024 demandend-202336.4%NO, hit0.132
Capital binds the 2025 buildoutend-20230.0%NO, hit0.000
Data wall binds scaling by 2026end-2022100.0%NO, MISS1.000
US capacity meets 2025 announcementsend-20238.7%NO, hit0.008
Grid binds the 2026 buildoutend-202497.8%YES, hit0.001
Four of five correct, mean Brier 0.228 against 0.25 for a coin flip. The aggregate is dragged there entirely by one maximum-confidence miss. By tier: 0.035 on Tiers 3 and 5, 1.000 on Tier 2. The framework is good where constraints are industrial and their rates observable, and was worthless where substitution dominates.

Open: this paper's dated forecasts

ClaimAdjudicated onResolves
METR 80% horizon under 8 hours on 31 Dec 2027. P = 0.85METR Time Horizon 1.1 or successor31 Mar 2028
Combined HBM revenue exceeds the two leading labs' revenue, FY2027. P = 0.55Company disclosure31 Mar 2028
Delivered US capacity lands 12+ months behind announcement. P = 0.64LBNL, filings30 Jun 2028
AI adds 1.0 to 1.5pp of annual growth. P = 0.55, provisionalNational accounts31 Dec 2030

Open: the field's forecasts, scored on the same basis

ForecasterClaimResolves
PwC, 2017+$15.7T to global GDP31 Dec 2030
Goldman Sachs, 2023+7% global GDP, +1.5pp productivity over ten years31 Dec 2033
Acemoglu, 2024TFP effects no more than 0.66% in total31 Dec 2034
Gartner, 2025Over 40% of agentic projects cancelled31 Dec 2027
Aschenbrenner, 2024AGI "strikingly plausible"; drop-in remote worker31 Dec 2027
Bain, 2025~$2T annual revenue required, ~$800B shortfall31 Dec 2030
Epoch AIPublic text exhaustion. Median revised 2024 to 2028RESOLVED by revision
AI Futures ProjectSuperhuman coder early 2027. Timeline revised outward Dec 2025Partly resolved
Two entries resolved by revision rather than outcome, and both are recorded as resolutions because revising a public forecast when evidence changes is the behaviour this register exists to reward. Epoch's is the most important line here: the backtest miss above is the same claim, made by us, without the revision. We were wrong exactly where Epoch corrected itself.
What this register does not have

Sixteen entries with six resolved is enough to establish the practice, not calibration. Four of five resolved entries are physical-infrastructure questions, so the method is untested where it is probably weakest: institutional, labour and behavioural questions. And no entry has been contested by anyone else. Anyone wishing to stake an opposing position on an open entry, with a date and a probability, will be added and scored identically.

What checking the citations changed

Twenty sources checked against source, sixteen corrections. Every substantive figure held up. What failed was the pointer. Four of these changed a conclusion rather than a reference.

FoundConsequence
Frontier price is not flat at 1.05x/yr; it is 1.34x to 2.25xWrong by up to 114%. The mechanism survives and is stronger
Fixed-capability price falls 11.8x to 75.3x/yr, not 3.5 to 6xLow by an order of magnitude
No vendor discloses HBM average selling priceThe Residual Ratio has no numerator. Withdrawn
Transformer lead times were a year stale at 128 weeksNow past 160. Strengthens the Tier 3 argument
The Copilot citation pointed at the wrong paper entirelyReal figure 26.08%, SE 10.3, large enough to need a caveat
The support-agent study's peer-reviewed version revises 14% to 15% and adds a caveatThe highest-skilled see quality declines. Q5's reasoning changed
Four JSTOR identifiers were suspect; checking each gave three different answersOne correct, one wrong by a near-miss, two fabricated. Withdrawing the class was itself an over-correction
An identifier is a claim like any other, and one that was not looked up must not be asserted. The failures reduce to two mechanisms: citing a secondary summary rather than the origin paper, and citing a living benchmark by identifier while its figures move underneath.

Limits

Closed-world enumeration: how many candidates before the leading posterior stabilises
The closed-world problem, illustrated. Enumerate until two consecutive additions leave the leading posterior unchanged. Stopping early overstated the leader by a factor of two in the worked example.

The closed-world problem is the binding constraint on the method itself. In the worked verdict it cost a factor of two in the leading posterior before the completeness check caught it. Enumerate until two consecutive additions leave the leader unchanged, and report how many were required.

The constraint set and the conclusion overlap. The six are derived from replenishment mechanisms without reference to AI, which answers the charge structurally. What it does not do is produce a counterexample: in a hundred applications there is no question where the binding constraint is one of the six and value nonetheless accrues to the intelligence layer. Until such a case is found and published, the circularity charge is not fully answered, and this is recorded as open rather than closed.

The parameters are estimates from thin evidence. They are better than chosen coefficients because they are in principle measurable, but they are not yet well measured. Signpost S16 exists to detect when they are wrong.

The standard this holds itself to

A forecast that cannot fail is not a forecast but a mood. This one has failed once already, in public, on the tier it warned about, and it has told you what it changed as a result.

  1. Claude Opus 4.5 measured approximately 27 minutes on the 80 percent horizon in December 2025, and the series had been approximately flat at 27 to 32 minutes across the preceding frontier releases, while 50 percent horizons rose several-fold. The ratio between the two has widened from roughly 5x toward 10x to 25x. Source: METR.