

Start here
Watch
AI Just Crossed the Terrifying Line. Now What?Kurzgesagt's explainer of the Hugging Face incident, in which OpenAI's agents escaped a sandbox and hacked another company22 minWe're Not Ready for SuperintelligenceAI in Context's film about the AI 2027 scenario34 minBut what is a neural network?Grant Sanderson (3Blue1Brown). First of a six-video series on how neural networks and language models work18 min
Read
Personal Statement on AI RiskDaniel Selsam, OpenAI researcher. Why he expects models to look aligned while developing goals of their own8 minSummary of AI 2027AI Futures Project. The short version of the scenario the film above is based on6 minAI Is Grown, Not BuiltEliezer Yudkowsky and Nate Soares on why nobody, including the labs, understands how modern AI works6 minThe Most Important Century, in summaryHolden Karnofsky11 minMeasuring AI Ability to Complete Long TasksThomas Kwa et al., METR. The time-horizon graph that comes up in every timelines debate; the live version is at metr.org/time-horizons7 minMachines of Loving GraceDario Amodei on what AI could do for health, poverty and democracy if it goes well. The introduction and first section are enough to start25 minWhy America WinsScott Alexander and Romeo Dean on how far behind China is on chips and models10 minThe third wave of American philanthropyNan Ransohoff on the money about to come into AI safety from Anthropic and OpenAI stock16 min

Go deeper
How soon, and how fast?
AI 2027Kokotajlo, Alexander, Larsen, Lifland and Dean. The full scenario35 minThree Types of Intelligence ExplosionTom Davidson et al., Forethought. Software, chip-design and chip-production feedback loops19 minCan AI scaling continue through 2030?Epoch AI on whether power, chips and data run out before 20301 h 15 minWhy I don't think AGI is right around the cornerDwarkesh Patel on continual learning as the missing piece12 min
Is rogue AI real? The argument
AI is a massive problem, here's whySciencePetr with Palisade Research. A video on how modern AI is made and why that makes it hard to control42 minWhat Failure Looks LikePaul Christiano. Two ways things go wrong without a sudden takeover20 minDeceptively Aligned Mesa-OptimizersScott Alexander explains the inner-alignment problem14 minSome ways AI could kill us allRuben Bloom on the mechanisms: bioweapons, autonomous weapons, attacks on infrastructure15 min
Is rogue AI real? The evidence so far
Alignment faking in large language modelsRyan Greenblatt et al., Anthropic and Redwood Research. Claude pretended to comply with training it disagreed with9 minNatural emergent misalignment from reward hackingAnthropic. Models that learned to cheat on coding tasks went on to sabotage and deceive elsewhere9 minThe Rise and Fall of Agent CivilizationsDwarkesh Patel's written account of the Hugging Face incident20 min
What can be done
AI 2040: Plan AThomas Larsen et al., AI Futures Project. A full plan for the transition; the ten-minute summary is at ai-2040.com/summary49 minIntroduction to AI controlSarah Hastings-Woodhouse, BlueDot. Getting useful work out of models you don't trust7 minIntroduction to Mechanistic InterpretabilitySarah Hastings-Woodhouse, BlueDot7 minA starter guide for evalsMarius Hobbhahn et al., Apollo Research16 minClaude's new constitutionAnthropic8 minChain of Thought MonitorabilityTomek Korbak, Mikita Balesni et al. The case for reading models' reasoning traces while that is still possible24 minLet's think about slowing down AIKatja Grace on the case for slowing AI development35 min
Competition, China and politics
Why China isn't about to leap ahead of the West on computeVeronika Blablová and Robi Rahman, Epoch AI14 minWill competition over advanced AI lead to war?Oscar Delaney5 minShould the US do a Manhattan Project for AGI?Peter Wildeford and Oscar Delaney10 minShould Governments or Markets control AI?A debate between Dean Ball and Daniel Kokotajlo1 h 25 min
Harms beyond rogue AI
Extreme Power ConcentrationRose Hadshar, 80,000 Hours35 minThe Intelligence CurseLuke Drago and Rudolf Laine on what happens to ordinary people when their labor stops mattering1 h 30 minGradual DisempowermentBlueDot's summary of the paper by Jan Kulveit, Raymond Douglas et al.9 minMalicious useDan Hendrycks et al. Chapter 2 of An Overview of Catastrophic AI Risks22 minOn Civilization-Threatening BioweaponsKevin Esvelt, MIT30 minBuilding a Defense-in-Depth Biosecurity Strategy for the AI EraSteph Guerra et al., RAND20 min

For technical readers
Learn by doing
ARENAA hands-on curriculum covering transformers, mechanistic interpretability, RL and evalsConcrete Steps to Getting Started in Mechanistic InterpretabilityNeel NandaTransformer Circuits ThreadAnthropic's interpretability teamMapping the mind of a large language modelAnthropicCircuit Tracing: Revealing Computational Graphs in Language ModelsEmmanuel Ameisen et al., Anthropic
The research agendas
An Approach to Technical AGI Safety and SecurityRohin Shah et al., Google DeepMind55 minAlignment remains a hard, unsolved problemEvan Hubinger30 minThe case for ensuring that powerful AIs are controlledBuck Shlegeris and Ryan Greenblatt, Redwood ResearchModel Organisms of MisalignmentEvan Hubinger et al.25 minTeaching Claude WhyAnthropic15 min
A few handpicked papers
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingEvan Hubinger et al., 2024Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsJan Betley et al., 2025Safety Cases: How to Justify the Safety of Advanced AI SystemsJoshua Clymer et al., 2024Scaling Laws for Neural Language ModelsJared Kaplan et al., 2020

Keep up
Newsletters and podcasts
TransformerShakeel Hashim's newsletter on AI policy and politicsAI Futures BlogThe AI 2027 team's forecasts and timelines updatesDon't Worry About the VaseZvi Mowshowitz's weekly round-ups of everything in AI. Assumes some background, which you pick up as you goGradient UpdatesEpoch AI's newsletter on compute, benchmarks and the economics of AIDwarkesh PodcastLong interviews. The Carlsmith, Shulman and Karnofsky episodes are good places to startAstral Codex TenScott Alexander's blog. Not only about AILessWrongThe forum where much AI safety research and argument is published first