Probability scoring ledger#
Contents
- §1 Round 0 - initial table
- §2 Round 8 - first re-score against rounds 1–7
- §3 Round 12 - commentary pass (no number moves)
- §4 Round 18 - re-score pass against rounds 13–17 (no number moves) + register
- §5 Round 25 - first grounded contradiction logged (no number moves)
- §6 Round 26 - re-score against the r25 Ground evidence (one move)
- §7 Rounds 27–35 - hold passes (no number moves)
- §8 How to read a hold
- §9 Resolution log (outcomes)
Per scoring rule 5: probabilities that never move are being defended, not updated. Every re-score of the Part V table is logged here - claim, prior, posterior, and the written reason. Mechanism accuracy matters more than outcome luck.
Round 0 - initial table#
| Row | Claim (short) | 2030 | 2040 | Note |
|---|---|---|---|---|
| 1 | Autonomous knowledge work | 45% | 80% | Initial |
| 2 | US TFP >2%/yr sustained | 25% | 55% | Initial |
| 3 | Major AI-attributed catastrophe | 20% | 45% | Initial |
| 4 | Binding international agreement + verification | 10% | 35% | Initial |
| 5 | Humanoids >1M units/yr deployed | 15% | 60% | Initial |
| 6 | AI-sector correction >40% | 40% | 65% | Initial |
| 7 | Existential loss of control | 1–3% | 3–8% | Initial |
Source: initial corpus, 2026-07-30. No prior.
Round 8 - first re-score against rounds 1–7#
Date: 2026-07-30. Corpus state: 69 files post–round 7, before round 8 expansions. Full rationale: reasoning.
| Row | 2030 prior → posterior | 2040 prior → posterior | Reason (one line) |
|---|---|---|---|
| 1 | 45% → 50% | 80% → 80% | Outcome pricing + junior-hiring data earlier than expected |
| 2 | 25% → 20% | 55% → 50% | Measured TFP is a high bar once quality-adjustment gap is priced (prices) |
| 3 | 20% → 20% | 45% → 45% | Mechanism detail added; no directional shift in severity |
| 4 | 10% → 8% | 35% → 30% | Compute-governance trap (bipolar): lever and verification are the same object |
| 5 | 15% → 12% | 60% → 55% | Robotics four-constraint split; unstructured path harder |
| 6 | 40% → 45% | 65% → 65% | Financing-mix migration raises credit-event odds (capital) |
| 7 | 1–3% → 1–3% | 3–8% → 3–8% | Three RSI governors bound the tail without rewriting the range |
Rows held with explicit justification (not inertia): 3 and 7. Rule 5 requires the justification when evidence existed and the number did not move; both are above.
Net directional read of the re-score: slightly more confident in near-term commercial and financial stress (rows 1, 6); slightly less confident in measured macro productivity, verified international control, and early humanoid scale (rows 2, 4, 5). The ordering thesis - incident >> loss-of-control as operative risk - is unchanged.
Round 12 - commentary pass (no number moves)#
Date: 2026-07-30. Corpus: post–rounds 9–11 (medicine, indicators, Game 2, winters) + r12 cyber/bio/education depth.
| Row | Action | Note |
|---|---|---|
| 3 | Hold | Game 2 multi-mechanism table + cyber trough + bio chain analysis improve which incident, not P(incident above threshold) enough to move 20/45 |
| 4 | Hold | Trap + "rules ≠ verification" already priced in r8 delta; r11 capture caveat affects quality of rules, not existence probability enough to respecify 8/30 |
| 1–2, 5–7 | Hold | No new material forcing a move; education/meaning are welfare channels outside these rows |
Rule 5 satisfied: evidence reviewed, holds justified in writing, not inertia.
Round 18 - re-score pass against rounds 13–17 (no number moves) + register#
Date: 2026-07-30. Corpus: post–rounds 13–17 (law/finance/science, media/software, ag/logistics, C8 provenance, compute/data, B12).
| Row | Action | Note |
|---|---|---|
| 1 | Hold | r13–14 law/finance/software depth refines which tasks verify cheaply, not the share of a knowledge worker's day that does; the verification ordering was already priced in r8 |
| 2 | Hold | r16–17 compute/data consistency work changes no macro input; B2's "silence is the base case" stands |
| 5 | Hold | B12 operationalizes the structured-first claim into triggers; an indicator getting sharper is not evidence about the outcome |
| 3–4, 6–7 | Hold | No new material in r13–17 bearing on incidents, agreements, corrections, or control |
Also this round: the distributed predictions register created - 16 probability-stamped claims outside this table are now indexed and scoreable, closing the invariant-4 gap. No register claim required a number move on the main table.
Rule 5 satisfied: evidence reviewed, holds justified in writing, not inertia.
Round 25 - first grounded contradiction logged (no number moves)#
Date: 2026-07-30. Source: second live Ground pass.
The open-weight lag, carried since authoring at ~9–15 months and embedded in the Game 1 prediction ("gap stable rather than collapsing"), measures ~3–6 months as of Jan–May 2026 (Epoch AI). This is the corpus's first live-source contradiction of a stamped claim component. No main-table row carries this figure directly, so no number moves; the affected prediction scores as written at resolution, with the ground note at source. Direction of the miss: faster diffusion - which presses row 3's incident-surface reasoning and Game 2's leaky-bucket result toward more leakage, not less, and weakens any future argument that frontier control chokepoints are sufficient policy levers.
Rule 5 satisfied: evidence reviewed, contradiction logged rather than absorbed silently.
Round 26 - re-score against the r25 Ground evidence (one move)#
Date: 2026-07-30. Trigger: the r25 grounded contradiction (open-weight lag ~3–6 months, Epoch AI) plus the refreshed A-family prints.
| Row | Action | Note |
|---|---|---|
| 1 | Hold | Capex and revenue prints say the build continues; nothing new on task-hour substitution or outcome-pricing share, which are this row's actual variables |
| 2 | Hold | Larger capex does not bear on measured TFP clearing 2%; the J-curve position is unchanged |
| 3 | Hold | The shortened lag widens the incident surface directionally, but one measured series with no deployment-severity datapoint is below the bar for moving a catastrophe probability; the mechanism list stands |
| 4 | −1 / −2 → 7% / 28% | The lag compression is direct evidence on this row's central object: verification of frontier training covers a shrinking capability share when near-frontier weights circulate ~3–6 months behind. Same direction as r8's export-control argument, now measured |
| 5 | Hold | IFR 54% and rare-earth concentration sharpen the supply-chain picture without evidence on unstructured autonomy or delivered $/hour |
| 6 | Hold | The dollar gap widened (capex ~$700–725B guided vs ~$55B model-layer revenue) but revenue roughly doubled y/y, so the ratio - the variable that matters - is ambiguous; A3's financing mix went unverified this pass. Watch A2 through 2027 before moving |
| 7 | Hold | No new evidence on cycle-time compression or eval integrity; the governors' status is unchanged |
Rule 5 satisfied: one move with written mechanism, six holds with named evidence. First table move since round 8.
Rounds 27–35 - hold passes (no number moves)#
Date: 2026-07-31. Depth rounds through r35; no new scoreable series crossed a Part V bar.
All seven rows held. Mechanism prose and indicators improved (Taiwan gray-zone table, Game 3 scoring cross-section, robotics fully-loaded cost, partial-clause rules, capex sync already reflected in r26 row-6 hold). Capex guidance ~$700–725B reconfirmed at source - does not force a row-6 move while the ratio remains ambiguous. Open-weight lag already priced into row 4 in r26. Rule 5: multi-round holds with named reasons, not silence.
How to read a hold#
The ledger's hold entries are its least glamorous and most informative content, because the failure modes of a subjective table are asymmetric in a specific way. The pressure on a maintained forecast runs toward motion - every round deposits new material, and moving a number is how an analyst demonstrates responsiveness - so a table that drifts a few points every pass is usually tracking the author's reading list, not the world. Rule 5 cuts the other way against anchoring: a number that never moves is being defended. The discipline that resolves the tension is the one visible above - every hold names the evidence reviewed and says why it was insufficient, which converts "no change" from an absence into a claim that can itself be wrong. A reader auditing this ledger should look for two smells: consecutive holds citing evidence that plainly bears on a row (anchoring wearing rule-5 clothing), and moves justified by material that merely restates the argument that set the prior (motion mistaken for updating). Rounds 12 and 18 are the pattern to hold future passes to: mechanism detail sharpened, numbers explicitly not moved, reasons in writing.
Note also what the ledger cannot show yet: whether the holds were right. Until claims resolve, this page measures process discipline, not accuracy - a perfectly maintained ledger of badly calibrated numbers would look identical to this one. That is not a flaw to fix; it is the reason the resolution log below exists and the reason it is empty.
Resolution log (outcomes)#
Empty until claims come due. First wave: 2027–28. See scoring "first claims due."
| Due | Claim | Source | Outcome | Framework correction |
|---|---|---|---|---|
| - | - | - | - | - |