Change any one to replay from that year. Each year keeps its dice, so only your choice differs.
Policy levers
Held fixed 2027–2040 across every simulated future.
Unknowns about the world
Each future draws these from a bell curve: the center you set, ± the spread.
Risk this year
Hover a choice to preview its effect.
100%
expected chance of getting this far without catastrophe
Your policy levers
2,000 simulated futures
Copies these settings, the formulas, the model's code and the results as one prompt. Paste it into Claude or another AI and ask it to poke holes in the assumptions.
Each line in the 3D view is one simulated future, 2027 → its outcome. Height is frontier capability; depth is alignment minus capability (positive means alignment is ahead). Futures that end in catastrophe fall to the floor in the year it strikes; dots mark where each future ends. Hover or tap a line to read its values.
Drag to orbit · scroll to zoom · x = year · y = capability · z = alignment − capability
Copy page
Copy this text
Your browser blocked automatic copying. The text is selected below. Press Ctrl+C or ⌘C, then paste it into an AI assistant.
2027 · The decade that decides
GAME OF AGI
You chair a newly formed council that steers frontier AI policy for the US-aligned bloc. Over the next 13 years, systems will likely become smarter than any human. Every year you face a decision about racing China, releasing open weights, biodefense, neural interfaces and alignment.
Each year the dice are rolled against three catastrophes: an engineered pandemic, loss of control, and great-power war. Reach the superintelligence threshold with alignment ready, or reach 2040 on a careful path.
Some things nobody knows yet: whether AGI turns out safe by default, how hard alignment really is, and whether insiders would blow the whistle on a power grab. Each game draws them in secret. You find out at the end.
This is a toy model built to make trade-offs tangible, not a forecast. Every probability comes from transparent, adjustable assumptions (see "How the model works"). Reasonable experts disagree about nearly all of them.
YOUR DECISIONS
Under the hood
How the model works
The world has ten state variables (0–100) and eight policy levers (0–1). Each year the simulator rolls once against three annual hazards, then advances the world. All code is in model.js, which is short enough to read in one sitting.
Timelines follow AI 2040. Capability 70 is an Automated Coder (AC), 90 an AI that dominates top human experts (TED-AI), 100 superintelligence (ASI). Under today's policies the model reaches AC around 2030 and ASI about a year later, matching AI 2040's takeoff forecast. AI labor speeds up AI research 3× at AC and ~40× at TED-AI. Effective speed scales with compute0.74 (10× less compute, 5.5× slower), and a hidden progress-speed factor spreads timelines out.
Pandemic risk ≈ 14% × [σ((open-model capability − 62)/6) + leakage from closed models] × (1 − biodefense)²
Loss of control ≈ (not safe by default) × 30% × σ((capability − alignment − 36)/6) × σ((capability − 78)/4) × (0.4 + 0.9·hidden internal AI)
War risk ≈ 5.5% × (1 − coordination)² × race closeness × strategic stakes
Superintelligence (capability = 100): goes well if safe by default, else with probability σ((alignment − bar)/6),
bar = 45 + 32·(hidden alignment difficulty) + 0.2·(hidden internal AI − 30) − 10·(BCI adoption)
Lock-in ≈ σ((power concentration − 64)/6) × (1 − 0.6·oversight) × (1 − whistleblower odds) × (1 − 0.6·plurality)
Research vs. weights. As in AI 2040, publishing research is separate from releasing weights. Research transparency spreads power (dozens of labs can follow the frontier), speeds up shared alignment work, pulls internal AI into view and builds trust, but it also helps China catch up. Open-weights releases spread power too, but they also put capable models into anyone's hands, which drives the bio risk.
Safety share. The safety lever is the share of AI compute and AI labor spent on safety, from 0% to 100%. Its value is linear up to 30% and then diminishes, because people and ideas become the bottleneck. At 100% the US bloc trains no new capabilities, which is a unilateral pause. China only partly slows down in response, and if China takes the lead its lab uses US alignment work only as far as coordination and published research allow.
Neural interfaces. BCIs help in three speculative ways. Augmented researchers do alignment work up to 1.5× faster. Augmented overseers can follow what thousands of AI copies are doing, which shrinks hidden internal AI. And humans who can check AI reasoning directly lower the alignment bar at the handoff by up to 10 points. Pointing AI labor at neurotech speeds BCIs up, but surgery, approval and rollout cap adoption at about 20 points a year. A BCI push without diplomacy concentrates power.
Hidden internal AI. AI 2040 argues most takeover risk comes from models labs run internally, before anyone outside can see them. Racing grows that gap, especially once AI automates R&D. Transparency and inspections shrink it. A large gap raises the yearly loss-of-control risk and the bar alignment must clear at the handoff.
The deal. When coordination is at least 60 and diplomacy is high, a verified deal holds frontier capability at top-expert level (78), like Plan A's 2035 pause. As in AI 2040 (about 48% collapse risk over ten years), a deal can collapse, about 6% a year, and stays dead once it does.
Hidden unknowns. Each future first draws four facts nobody knows yet, each from a bell curve (mean ± spread): whether AGI is safe by default, how difficult alignment really is, how fast AI progresses, and the odds that whistleblowers stop a power grab. The risk panel shows risks averaged over those unknowns. The dice use the hidden truth.
Calibration. Holding each plan's levers fixed, the model's chance of getting alignment right is about 18% for racing to ASI (AI 2040's median estimate: 25%), 40% for burning the lead (40%) and 79% for Plan A (72%).
What it leaves out. Covert projects, cyberattacks, other countries, misuse beyond bio, conflict between AGIs, and whether "capability" can really be put on one axis at all. Treat it as an argument you can poke at, not an oracle.
This version extends the original Game of AGI, a Monte Carlo of AGI risk driven by safe-by-default odds, safety spending, oversight, the number of labs that originate AGI, whistleblowers, and takeoff vs. replication time. It adds time, geopolitics, open weights and bio risk, and calibrates timelines to AI 2040.
How it was built. This game was built in conversation with Claude Code. Each session's full history is browsable as a memtree:
All parameters are illustrative. The lab, company and leaders in this game are fictional composites.