The Next Fifteen Years

A forecast built from first principles
Section future / 08-method / scoring.md

Scoring - how to hold this document accountable#


Contents

Predictions that cannot be scored are decoration. This page specifies how each claim resolves, what counts as a miss, and the specific ways this framework will try to avoid admitting one.

Resolution rules#

A blockquoted prediction resolves. Every > block in the corpus is a scoreable claim. Prose is context.

Ranges resolve at the midpoint unless a threshold is stated. "12–18% of task-hours" scores against 15%. "By 2029" means by 2029-12-31.

Directional claims resolve against the counterfactual, not the level. "Value accrues to inelastic complements" is not scored by whether energy prices rose - they might rise for unrelated reasons - but by whether they rose relative to cognition-intensive prices. Most of this document's claims are relative and must be scored relative.

Relative claims fix their comparator series at authoring. "Trades wages outpace credentialed professions" scores against the series pair named where the claim lives, not against whichever wage series later flatters it. Comparator shopping is reference-class swapping wearing a smaller hat, and it is the specific dodge relative claims invite.

A claim is a miss if it was right for the wrong reason. If the correction arrives in 2028 but is triggered by a macro shock rather than by AI revenue disappointing, the timing was lucky and the mechanism was wrong. Log it as a miss with a note. Mechanism accuracy is the point; a framework that gets outcomes right through wrong mechanisms will fail on the next question.

The scoring ledger#

Maintained per protocol §5. Each entry: claim, source page, resolution date, outcome, and - the important column - what the framework should have said.

FieldWhy
Claim, verbatimPrevents retroactive rewording, the most common self-deception
Source pageLocates the reasoning that produced it
Resolution date and criterionFixed at authoring time, never after
OutcomeHit / miss / right-for-wrong-reason / unresolvable
Framework correctionThe only field that produces learning

Unresolvable is a real and common outcome, and it is a mild indictment. A claim that cannot be settled by evidence was underspecified when written. Count them; a rising unresolvable rate means the writing is drifting toward safety.

Who scores, and the conflict in it#

The document scores itself, which is a conflict of interest with a known failure profile: leniency on ambiguous resolutions and creative readings of "the mechanism was right." Two mitigations, neither complete. First, the verbatim-claim rule exists precisely so an outside reader can re-score from the ledger without trusting the scorer - the ledger is designed to be auditable, not just kept. Second, ambiguous resolutions default to miss. A rule where ties go to the house selects for vague writing; a rule where ties go against it selects for sharper claims in the next round, which is the incentive the whole exercise is trying to create. The residual weakness is claim selection rather than claim scoring - the temptation to blockquote only the safe claims - and the unresolvable-rate count above is the partial check on it.

The five ways this document will dodge#

Named in advance, because the point of naming them is to make them harder to use.

  1. "Early, not wrong." The universal escape. Rule: a timing claim that misses its stated date by more than 3 years is a miss, regardless of what happens afterward. Directional vindication at an arbitrary horizon is not a forecast.
  2. Reference-class swapping. When aviation stops fitting the regulatory prediction, reaching for privacy instead. Rule: the class is fixed at authoring. Changing it is a logged revision with a stated reason, not an interpretation.
  3. Scope narrowing. Game 3's narrowing to "competitive markets specifically" in the steelman is exactly this move. It is legitimate once, done explicitly, in advance of the resolution. It is not legitimate afterward.
  4. Indicator substitution. Quietly promoting whichever indicator happens to be confirming. Rule: the five headline indicators are fixed; changes go in the round log.
  5. Probability inertia. The subtlest. A probability that never moves is not being updated; it is being defended. Rule: annual re-scoring of every row in Part V, and a row that hasn't moved in two years requires a written justification for why no evidence bore on it.

Small-sample honesty#

The calibration table below assumes enough resolved claims for frequencies to mean something, and for years it will not have them. With a few dozen resolutions, a perfectly calibrated 60% band still produces runs that look like systematic bias, and an ill-calibrated one can look fine. Until the ledger is large, treat calibration statistics as suggestive and mechanism post-mortems as the real signal - a single right-for-the-wrong-reason resolution teaches more at small N than the hit rate does. The failure mode this guards against is premature vindication: declaring the framework calibrated off a sample that could not have shown otherwise.

Calibration targets#

Over a sufficient number of resolved claims:

Stated confidenceShould be right
~90%~9 in 10
~60%~6 in 10
~20%~2 in 10

The 60% band is where this document lives and where calibration is hardest to fake. Being right on 90% of the 60% claims is not good performance - it is evidence of systematic under-confidence, which is its own failure and usually a sign of hedging to avoid scoreable error.

The first claims due#

DueClaimSource
2027–28Junior hiring fails to recover with aggregate white-collar hiringGame 4, B1
2028No frontier training run above $5B → the capex ceiling was evadableA1
2028Revenue above ~$200B/yr run-rateA2
2027–2031A salient AI-attributed incident occursGame 2, C1
2029Consolidation to 3–5 frontier labs; open-weight tier 9–15 months behindGame 1
2029>40% of new US frontier capacity on owned or bilaterally-contracted generationEnergy
2030AI-specific liability coverage exists as a material line with published ratesInsurance

The first two resolve within about eighteen months of writing and are the earliest real test. If both miss, the framework's treatment of institutional friction and capital ceilings is wrong in the same direction - toward slowness - and the whole timeline should shift, not just those two rows.

Partial credit and the lag clause#

Game 1's consolidation prediction bundles two clauses: lab count by 2029, and open-weight lag ~9–15 months stable. The lag clause is already running wrong (measured ~3–6 months). Score by clause, not by the conjunction: a correct consolidation count with a wrong lag is a partial hit that still revises Game 2 leakage and row 4, which is what r26 did. Refusing to score until 2029 would bury an already-resolved component - one of the five dodges this page exists to prevent.

Evidence family by period (from timelines hub): 2026–28 financial; 2028–32 statistical/legal; 2032–40 physical. Scoring a claim with the wrong family is a mechanism error even if the calendar outcome lucks out.


Related: Part V - Probabilities · Part VII - Indicators · Steelman · Dependencies · Notation · Protocol

View markdown source

select · Enter open · Esc close