The Next Fifteen Years

A forecast built from first principles
Section future / 05-probabilities / ledger.md

Probability scoring ledger#


Contents

Per scoring rule 5: probabilities that never move are being defended, not updated. Every re-score of the Part V table is logged here - claim, prior, posterior, and the written reason. Mechanism accuracy matters more than outcome luck.

Round 0 - initial table#

RowClaim (short)20302040Note
1Autonomous knowledge work45%80%Initial
2US TFP >2%/yr sustained25%55%Initial
3Major AI-attributed catastrophe20%45%Initial
4Binding international agreement + verification10%35%Initial
5Humanoids >1M units/yr deployed15%60%Initial
6AI-sector correction >40%40%65%Initial
7Existential loss of control1–3%3–8%Initial

Source: initial corpus, 2026-07-30. No prior.

Round 8 - first re-score against rounds 1–7#

Date: 2026-07-30. Corpus state: 69 files post–round 7, before round 8 expansions. Full rationale: reasoning.

Row2030 prior → posterior2040 prior → posteriorReason (one line)
145% → 50%80% → 80%Outcome pricing + junior-hiring data earlier than expected
225% → 20%55% → 50%Measured TFP is a high bar once quality-adjustment gap is priced (prices)
320% → 20%45% → 45%Mechanism detail added; no directional shift in severity
410% → 8%35% → 30%Compute-governance trap (bipolar): lever and verification are the same object
515% → 12%60% → 55%Robotics four-constraint split; unstructured path harder
640% → 45%65% → 65%Financing-mix migration raises credit-event odds (capital)
71–3% → 1–3%3–8% → 3–8%Three RSI governors bound the tail without rewriting the range

Rows held with explicit justification (not inertia): 3 and 7. Rule 5 requires the justification when evidence existed and the number did not move; both are above.

Net directional read of the re-score: slightly more confident in near-term commercial and financial stress (rows 1, 6); slightly less confident in measured macro productivity, verified international control, and early humanoid scale (rows 2, 4, 5). The ordering thesis - incident >> loss-of-control as operative risk - is unchanged.

Round 12 - commentary pass (no number moves)#

Date: 2026-07-30. Corpus: post–rounds 9–11 (medicine, indicators, Game 2, winters) + r12 cyber/bio/education depth.

RowActionNote
3HoldGame 2 multi-mechanism table + cyber trough + bio chain analysis improve which incident, not P(incident above threshold) enough to move 20/45
4HoldTrap + "rules ≠ verification" already priced in r8 delta; r11 capture caveat affects quality of rules, not existence probability enough to respecify 8/30
1–2, 5–7HoldNo new material forcing a move; education/meaning are welfare channels outside these rows

Rule 5 satisfied: evidence reviewed, holds justified in writing, not inertia.

Round 18 - re-score pass against rounds 13–17 (no number moves) + register#

Date: 2026-07-30. Corpus: post–rounds 13–17 (law/finance/science, media/software, ag/logistics, C8 provenance, compute/data, B12).

RowActionNote
1Holdr13–14 law/finance/software depth refines which tasks verify cheaply, not the share of a knowledge worker's day that does; the verification ordering was already priced in r8
2Holdr16–17 compute/data consistency work changes no macro input; B2's "silence is the base case" stands
5HoldB12 operationalizes the structured-first claim into triggers; an indicator getting sharper is not evidence about the outcome
3–4, 6–7HoldNo new material in r13–17 bearing on incidents, agreements, corrections, or control

Also this round: the distributed predictions register created - 16 probability-stamped claims outside this table are now indexed and scoreable, closing the invariant-4 gap. No register claim required a number move on the main table.

Rule 5 satisfied: evidence reviewed, holds justified in writing, not inertia.

Round 25 - first grounded contradiction logged (no number moves)#

Date: 2026-07-30. Source: second live Ground pass.

The open-weight lag, carried since authoring at ~9–15 months and embedded in the Game 1 prediction ("gap stable rather than collapsing"), measures ~3–6 months as of Jan–May 2026 (Epoch AI). This is the corpus's first live-source contradiction of a stamped claim component. No main-table row carries this figure directly, so no number moves; the affected prediction scores as written at resolution, with the ground note at source. Direction of the miss: faster diffusion - which presses row 3's incident-surface reasoning and Game 2's leaky-bucket result toward more leakage, not less, and weakens any future argument that frontier control chokepoints are sufficient policy levers.

Rule 5 satisfied: evidence reviewed, contradiction logged rather than absorbed silently.

Round 26 - re-score against the r25 Ground evidence (one move)#

Date: 2026-07-30. Trigger: the r25 grounded contradiction (open-weight lag ~3–6 months, Epoch AI) plus the refreshed A-family prints.

RowActionNote
1HoldCapex and revenue prints say the build continues; nothing new on task-hour substitution or outcome-pricing share, which are this row's actual variables
2HoldLarger capex does not bear on measured TFP clearing 2%; the J-curve position is unchanged
3HoldThe shortened lag widens the incident surface directionally, but one measured series with no deployment-severity datapoint is below the bar for moving a catastrophe probability; the mechanism list stands
4−1 / −2 → 7% / 28%The lag compression is direct evidence on this row's central object: verification of frontier training covers a shrinking capability share when near-frontier weights circulate ~3–6 months behind. Same direction as r8's export-control argument, now measured
5HoldIFR 54% and rare-earth concentration sharpen the supply-chain picture without evidence on unstructured autonomy or delivered $/hour
6HoldThe dollar gap widened (capex ~$700–725B guided vs ~$55B model-layer revenue) but revenue roughly doubled y/y, so the ratio - the variable that matters - is ambiguous; A3's financing mix went unverified this pass. Watch A2 through 2027 before moving
7HoldNo new evidence on cycle-time compression or eval integrity; the governors' status is unchanged

Rule 5 satisfied: one move with written mechanism, six holds with named evidence. First table move since round 8.

Rounds 27–35 - hold passes (no number moves)#

Date: 2026-07-31. Depth rounds through r35; no new scoreable series crossed a Part V bar.

All seven rows held. Mechanism prose and indicators improved (Taiwan gray-zone table, Game 3 scoring cross-section, robotics fully-loaded cost, partial-clause rules, capex sync already reflected in r26 row-6 hold). Capex guidance ~$700–725B reconfirmed at source - does not force a row-6 move while the ratio remains ambiguous. Open-weight lag already priced into row 4 in r26. Rule 5: multi-round holds with named reasons, not silence.

How to read a hold#

The ledger's hold entries are its least glamorous and most informative content, because the failure modes of a subjective table are asymmetric in a specific way. The pressure on a maintained forecast runs toward motion - every round deposits new material, and moving a number is how an analyst demonstrates responsiveness - so a table that drifts a few points every pass is usually tracking the author's reading list, not the world. Rule 5 cuts the other way against anchoring: a number that never moves is being defended. The discipline that resolves the tension is the one visible above - every hold names the evidence reviewed and says why it was insufficient, which converts "no change" from an absence into a claim that can itself be wrong. A reader auditing this ledger should look for two smells: consecutive holds citing evidence that plainly bears on a row (anchoring wearing rule-5 clothing), and moves justified by material that merely restates the argument that set the prior (motion mistaken for updating). Rounds 12 and 18 are the pattern to hold future passes to: mechanism detail sharpened, numbers explicitly not moved, reasons in writing.

Note also what the ledger cannot show yet: whether the holds were right. Until claims resolve, this page measures process discipline, not accuracy - a perfectly maintained ledger of badly calibrated numbers would look identical to this one. That is not a flaw to fix; it is the reason the resolution log below exists and the reason it is empty.

Resolution log (outcomes)#

Empty until claims come due. First wave: 2027–28. See scoring "first claims due."

DueClaimSourceOutcomeFramework correction
-----

Related: Reasoning · Scoring · Protocol §5

View markdown source

select · Enter open · Esc close