Science#
Contents
Highest-value application, and the most likely source of genuine compounding growth. If anything in this document produces a permanent change in the level of human welfare rather than a redistribution, it is this.
The verification asymmetry, applied#
Progress will be badly uneven, in a predictable order set by the cost of ground truth:
- Math and theoretical CS - fastest. Proof checkers give free, perfect, instant verification.
- Computational chemistry and materials - fast. Simulation is imperfect but cheap.
- Experimental biology - slow. Ground truth requires a wet lab and months.
- Social science - slowest. Ground truth is contested, expensive, and often unobtainable.
This is the data asymmetry expressed as a research agenda. Same ordering as Part III's domain table, applied inside research itself.
The actual bottleneck#
AI's contribution to hypothesis generation is already large. Its contribution is bottlenecked on experimental cycle time.
The constraint is no longer ideas. It is the rate at which reality can be asked questions.
Whoever industrializes automated experimentation - self-driving labs at scale - captures the largest available prize of the 2030s.
This is the most under-invested area in the corpus relative to what it would buy. It converts an expensive verification signal into a cheap one, which - per Part I - is precisely the move that lets capability grow in a domain. It is also unglamorous, capital-intensive, and physical, which is why it is underfunded relative to model work.
Selection becomes the scarce skill#
When hypotheses were expensive, generating a good one was the mark of a scientist. When they are nearly free, the binding skill inverts: deciding which of ten thousand plausible hypotheses deserves one of your finite lab-months. That is portfolio allocation under deep uncertainty - expected information gain per dollar of experiment - and it is trained by exactly the slow bench experience that automation displaces, the same apprenticeship loop as software review. Two consequences. First, groups with instrumented feedback on their own selection quality (did our chosen experiments outperform the ones we skipped?) will compound advantage the way test-covered codebases do; almost no lab currently measures this. Second, cheap hypothesis generation raises the value of negative results and shared failure data, because the cost of everyone independently testing the same seductive wrong idea scales with the generation rate. The current publication system discards negative results almost perfectly, which means the waste scales with model capability until the incentive is fixed. Failure mode: if learned selection models beat human taste at ranking experiments (a narrower, more checkable claim than general verification), the inversion favors whoever has the historical outcomes data, and incumbent pharma screening archives become one of the quietly valuable datasets on earth.
Automated labs: what they buy#
| Capability | Effect | Still binds |
|---|---|---|
| Closed-loop design–run–measure | Compresses cycle time where assays are automatable | Assay development, edge cases |
| Parallel cloud labs / CRO robotics | Throughput without every PI owning hardware | Queue price, standardization |
| Literature + protocol agents | Faster setup, fewer dumb failures | Wet-lab tacit skill (biosecurity dual use) |
| Simulation-in-the-loop | Fewer physical trials per hit | Sim-to-real gap |
Automated labs are the RSI physical twin: Uncertainty 1 needs validated experiments; science automation is how validation escapes human hands. The three governors still apply - verification of whether the science was right, physical supply chains for instruments and reagents, and financial hurdle rates on lab capex.
Drug discovery bridge#
Drug discovery: AI accelerates preclinical design; ~60% of clinical failures are human efficacy/toxicity; clinical gauntlet barely moves before ~2032 without better translational biology and trial logistics.
| Science automation helps | Does not replace |
|---|---|
| Target ID, design, in vitro loops | Phase II/III clocks, recruitment |
| Materials for delivery / manufacturing | Regulatory evidence standards |
| Biomarker and assay invention | Hospital and CRO capacity |
Estimate (aligned with medicine page): ~20–30% preclinical cost reduction early; clinical timeline compression is a lab + regulation story, not a model-release story.
Bio defense bridge#
Same stack, opposite sign of biosecurity:
| Defensive use | Why it matters |
|---|---|
| Surveillance + anomaly detection | Compresses detection; everything downstream scales with it |
| Countermeasure design | Design half already moving |
| Surge manufacturing science | The slow term - highest marginal defensive $ |
The gap remains physical and regulatory pipeline, not design cleverness. Automated labs without access controls also worsen the offense side (protocol + production compression). Dual-use is not a slogan here; it is the same equipment.
Screening (synthesis providers) stays the best non-lab control. Lab automation policy is the hard complement: who may run which closed loops on which agents.
Energy and materials#
Energy sector Layer 4: fusion, advanced fission, storage, catalysts - AI multiplies design and simulation; licensing and FOAK construction remain rate limits. Materials discovery is the cleanest automated-lab ROI outside biology when ground truth is instrumented.
Institutional absorption#
Science is not only labs - it is journals, grants, tenure, and IP:
- Publication signal collapses under paper mills and fluent text (Game 5); provenance and replication become the scarce goods
- Grant systems optimize for proposal fluency unless review becomes synchronous / work-sample based
- IP and data access determine who can close the loop (proprietary assays vs open science)
Universities face the same education bind: content cheap, assessment and lab seats expensive.
What to watch#
| Signal | Reading |
|---|---|
| Capex and utilization of automated / cloud labs | Prize being pursued |
| Time from hypothesis to validated result in instrumented fields | Real compression |
| Phase II success rates (pharma) | Translational wall |
| Synthesis screening coverage | Bio control point |
| Replication / provenance requirements at top venues | Signal repair |
Failure modes#
- If sim-to-real is good enough, physical lab bottleneck softens earlier than claimed.
- If automated labs stay boutique, the "largest prize" claim fails on capital allocation, not on idea quality.
- If learned verification (Uncertainty 5) makes theoretical fields self-validate wrongly, math/CS speed becomes a liability - confident wrong proofs at scale.
Paper flood is the free generation problem inside science#
The same verification scarcity that reorders capability also hits the literature: more papers, faster, with weaker average signal. Top venues raising provenance and replication bars are not Luddism - they are C8-style infrastructure for claims. Fields that can check cheaply (some ML, some formal math) will absorb the flood; fields that cannot will either slow acceptance or degrade. That is Game 5 with tenure on the line. Watch desk-reject rates and replication mandates before watching raw publication counts as "progress."
Automated labs are the 2030s prize, not another model release. Capex and utilization of cloud/robotic wet labs move this page; leaderboard wins on pure theory do not. Boutique demo labs with no utilization series are theatre. → drug discovery
Related: Data · Drug discovery · Biosecurity · Energy sector · Uncertainty 1 · 2032–2040