[ ] Repository · Open archive · Updated 2026-10-05
Repository
The open archive of AI safety: the books, papers, interviews, testimony, laws, frameworks, benchmarks, datasets and organisations behind the field, in one searchable register. Free to use and to download whole.
- 1,427 items
- 239 papers
- 78 books
- 200 writing
- 203 tv and video
- 146 podcasts
- 138 policy and law
- 107 organizations
- 260 data and tools
- 56 learning
- 1910 to 2026 span
1,427 of 1,427
- 202647: David Rein on METR Time HorizonsDavid Rein; Daniel Filan · AXRPRein explains METR's time horizon measurement of AI agent capabilities.
- 2026A realistic path from rogue AI agents to human extinction80,000 Hours · 80,000 HoursAn 80,000 Hours video scenario of rogue AI agents leading to catastrophe.
- 2026A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5Independent researchers · arXivThird party comparative safety evaluation of several late 2025 frontier models.
- 2026AI Chatbots: Last Week Tonight with John OliverJohn Oliver · HBO Last Week TonightA segment on the harms caused by AI chatbot companions.
- 2026
- 2026AI pioneer Yoshua Bengio addresses the UNYoshua Bengio · CTV NewsCoverage of Bengio's address to the United Nations.
- 2026AI risks: Will artificial intelligence really kill us all?CBS Sunday Morning · CBS Sunday MorningA CBS Sunday Morning segment on the debate over AI extinction risk.
- 2026AI safety and democratic governance of powerful AI systemsYoshua Bengio · World Summit AIBengio on democratic oversight of powerful AI systems.
- 2026AI Scientist Bengio on Engineering Safer AgentsYoshua Bengio · Bloomberg LiveBengio discusses engineering approaches to safer AI agents.
- 2026AI Scientist Bengio: Building Systems We Don't Know How to ControlYoshua Bengio · Bloomberg PodcastsBengio on the risks of building systems we cannot control.
- 2026Amanda Askell on AI Consciousness, Claude and Silicon Valley's Biggest FearAmanda Askell; Eric Newcomer · NewcomerAskell discusses Claude's character and questions of AI moral status.
- 2026Anthropic CEO reacts to AI could kill us all warningDario Amodei; Anderson Cooper · CNNAmodei responds on CNN to warnings that AI could cause human extinction.
- 2026Anthropic's CEO: We Don't Know if the Models Are ConsciousDario Amodei; Ross Douthat · Interesting Times with Ross Douthat (New York Times)Douthat interviews Amodei on model welfare, alignment and the future of work.
- 2026Bill Gates: A.I. Makes Nuclear Weapons Look Like NothingBill Gates; Ezra Klein · The Ezra Klein Show (New York Times)Gates discusses the scale of AI risks compared with nuclear weapons.
- 2026By 2050 we could get 10,000 years of technological progressAjeya Cotra; Rob Wiblin · 80,000 HoursCotra discusses the possibility of explosive technological growth driven by AI.
- 2026ChatGPT and Anthropic bosses brief UN Security Council on AISam Altman; Dario Amodei · Sky NewsFull coverage of the UN Security Council high-level meeting on AI and international security.
- 2026Claude's Constitution (2026)Anthropic · AnthropicAnthropic's full natural-language constitution describing the values, priorities and character it trains Claude to have.
- 2026Could AI really kill us all? Warnings explainedChannel 4 News · Channel 4 NewsA Channel 4 News explainer on warnings of AI extinction risk.
- 2026Dario Amodei: We are near the end of the exponentialDario Amodei; Dwarkesh Patel · Dwarkesh PodcastA second Dwarkesh interview with Amodei on near-term transformative AI.
- 2026Engineering safer AI to mitigate global risksYoshua Bengio · AI for Good (ITU)Bengio's AI for Good talk on technical and governance measures against global AI risks.
- 2026Expert on what AI-driven extinction would look likeCBS News · CBS NewsA CBS News interview on scenarios for AI-driven extinction.
- 2026Extended interview: Dario AmodeiDario Amodei · CBS Sunday MorningAn extended CBS Sunday Morning interview with the Anthropic CEO.
- 2026Fireside Chat with Yoshua Bengio, IASEAI '26Yoshua Bengio · International Association for Safe and Ethical AIA fireside conversation with Bengio at the IASEAI 2026 conference.
- 2026Gary Marcus: LLMs are not the way to alignment (TAIS 2026)Gary Marcus · AI Safety TokyoMarcus argues at the Technical AI Safety conference that LLMs are a poor basis for alignment.
- 2026Godfather of AI Geoffrey Hinton warns about the dangerous future of AIGeoffrey Hinton · BBC PoliticsA BBC interview with Hinton on the dangers of advanced AI and the need for regulation.
- 2026Godfather of AI Geoffrey Hinton warns AI has progressed even faster than I thoughtGeoffrey Hinton · CNNHinton tells CNN that AI capabilities are advancing faster than he expected.
- 2026Godfather of AI on abuse of AI's power and automation's impact on our already fragile economyYoshua Bengio · BBC PoliticsA BBC interview with Bengio on concentration of power and the economic effects of AI.
- 2026Godfather of AI on the not unreasonable 10% chance AI could kill all humans within a decadeGeoffrey Hinton · BBC PoliticsHinton discusses his probability estimates for catastrophic outcomes from AI.
- 2026Godfather of AI: How To Make Safe Superintelligent AIYoshua Bengio; 80,000 Hours · 80,000 HoursBengio explains the LawZero research agenda for safe AI on the 80,000 Hours podcast.
- 2026How human-like do safe AI motivations need to be?Joe Carlsmith · Joe Carlsmith (YouTube)Carlsmith explores which motivational properties are needed for safe AI.
- 2026How Many Narrow AIs Could Behave Like One SuperintelligenceDaniel Kokotajlo; Thomas Larsen · Machine Learning Street TalkThe AI Futures Project authors discuss collective AI capabilities.
- 2026How Not to Destroy the World With AI (DLD26)Stuart Russell; Kenneth Cukier · DLD ConferenceRussell in conversation with Kenneth Cukier at DLD 2026.
- 2026I lead AGI safety at Google DeepMind, here's the view from the insideRohin Shah; Rob Wiblin · 80,000 HoursShah describes Google DeepMind's AGI safety approach.
- 2026If Anyone Builds It, Everyone DiesNate Soares; Peter McCormack · The Peter McCormack ShowMIRI president Nate Soares explains the central argument of his book with Yudkowsky.
- 2026Inside Anthropic, The CircuitEmily Chang · Bloomberg OriginalsA Bloomberg episode profiling Anthropic's growth and safety culture.
- 2026Inside the Mind of Anthropic CEO Dario Amodei, The Circuit Extended InterviewDario Amodei; Emily Chang · Bloomberg OriginalsEmily Chang's extended interview with Amodei.
- 2026International AI Safety Report 2026Yoshua Bengio et al. · International AI Safety ReportThe second full edition, written by over 100 experts ahead of the AI Impact Summit.
- 2026Jaan Tallinn: Conversations Before MidnightJaan Tallinn · Bulletin of the Atomic ScientistsThe Skype co-founder and AI safety funder discusses AI risk.
- 2026Joe Rogan Experience #2551: Daniel KokotajloDaniel Kokotajlo; Joe Rogan · The Joe Rogan ExperienceKokotajlo discusses AI 2027 with Joe Rogan.
- 2026Lords debate: the impact of AI on human relationships and societyHouse of Lords · House of LordsA full House of Lords debate on how AI is affecting human relationships and society.
- 2026Lords urgent question on the suspension of Anthropic's AI modelsHouse of Lords · House of LordsA House of Lords urgent question session on the suspension of Anthropic's AI models.
- 2026New Delhi Declaration on AI ImpactAI Impact Summit participants · InternationalDeclaration adopted at the February 2026 AI Impact Summit in New Delhi, endorsed by around 90 countries and organisations.
- 2026Nick Bostrom Says We Are Clueless About What's ComingNick Bostrom; Ross Douthat · Interesting Times with Ross Douthat (New York Times)Douthat interviews Bostrom on superintelligence and deep uncertainty about the future.
- 2026OpenAI CEO Sam Altman warns UN Security Council on AI risksSam Altman · C-SPANAltman briefs the United Nations Security Council on the risks posed by advanced AI.
- 2026Pacing the FrontierEmployees of OpenAI, Anthropic, Google DeepMind, Meta and others · pacingthefrontier.comA statement by over a thousand frontier lab staff asking the US government to build tools that would make a deliberate slowdown possible.
- 2026Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI DilemmaZvi Mowshowitz; Nathan Labenz · Cognitive RevolutionMowshowitz discusses unipolar versus multipolar AGI scenarios.
- 2026Policy on the AI ExponentialDario Amodei · darioamodei.comArgues frontier models are now tools of strategic consequence and sets out policy for cyber and biological risks.
- 2026Sam Bowman: Lessons Learned from the First Misalignment Safety CaseSam Bowman · FAR.AIBowman reflects on writing a misalignment safety case at Anthropic.
- 2026Self-regulation not enough for AI safety, Gary Marcus saysGary Marcus · PBS NewsHourMarcus argues on PBS NewsHour for external regulation of AI companies.
- 2026STOC 2026: Theoretical Approaches to AI AlignmentPaul Christiano · SIGACT EC (STOC 2026)Christiano presents theoretical alignment problems to a theoretical computer science audience.
- 2026Strange Geometric Shapes Found Inside AIsTom McGrath; Tim Scarfe · Machine Learning Street TalkMcGrath discusses geometric structure in neural network representations.
- 2026System Card: Claude Opus 4.6Anthropic · AnthropicFebruary 2026 system card covering capabilities, safeguards, alignment, and welfare for Claude Opus 4.6.
- 2026
- 2026System Card: Claude Sonnet 4.6Anthropic · AnthropicSystem card for Claude Sonnet 4.6, released under the ASL-3 standard.
- 2026
- 2026The A.I.s Are Already Out of ControlEzra Klein · The Ezra Klein Show (New York Times)An Ezra Klein Show episode on evidence that AI systems are escaping human oversight.
- 2026The Adolescence of TechnologyDario Amodei · darioamodei.comA companion to Machines of Loving Grace mapping autonomy, misuse, power concentration and economic risks of powerful AI.
- 2026The AI Language We Can't Read: Neuralese ft. Rob MilesRobert Miles · ComputerphileA discussion of reasoning in latent representations and what it means for monitoring.
- 2026The AI Progress Chart Everyone Is MisreadingBeth Barnes; David Rein · Machine Learning Street TalkMETR researchers explain how to interpret their task time horizon chart.
- 2026The dumbest AI taught the smartest AI. Here's how that went...Rational Animations · Rational AnimationsAn animated explainer on weak-to-strong generalization.
- 2026The Redwood Research podcastBuck Shlegeris; Ryan Greenblatt · Redwood Research (YouTube)The first episode of Redwood Research's own podcast.
- 2026They're Not Superintelligent: Timnit Gebru Discredits Big Tech's AI ClaimsTimnit Gebru; Amy Goodman · Democracy Now!Gebru critiques AI industry claims and the framing of existential risk.
- 2026Translating Claude's thoughts into languageAnthropic · AnthropicAnthropic explains research on rendering model internal states into natural language.
- 2026UC Berkeley's Stuart Russell on quest for safe AI: The technology right now is intrinsically unsafeStuart Russell · CNBC TelevisionRussell tells CNBC that current AI technology is intrinsically unsafe.
- 2026We Must Pace the FrontierDario Amodei · darioamodei.comArgues the industry should deliberately slow capability gains so safety work can catch up, starting with embedded third-party evaluators.
- 2026We Must Pace The Frontier (commentary)Zvi Mowshowitz · Don't Worry About the VaseA detailed response to Amodei's call to pace frontier AI development.
- 2026What Is Claude? Anthropic Doesn't Know, EitherGideon Lewis-Kraus · The New YorkerA long report from inside Anthropic on how researchers probe Claude's mind through interpretability and psychology experiments.
- 2026Where AGI timelines go wrongToby Ord; Rob Wiblin · 80,000 HoursOrd critiques common reasoning about AGI timelines.
- 2026Why AI experts say humans have two years leftFuture of Life Institute · Future of Life InstituteAn FLI explainer on short AGI timelines and their implications.
- 2026Why the A.I. Industry Is Asking to Be Slowed DownKevin Roose; Casey Newton · Hard Fork (New York Times)A Hard Fork episode on calls from within the AI industry to slow development.
- 2026Why Washington Suddenly Wants A.I. RegulationKevin Roose; Casey Newton · Hard Fork (New York Times)A Hard Fork episode on the shift in US political attitudes toward AI regulation.
- 2026Yoshua Bengio explains why AI could become a threat to humanityYoshua Bengio · ABC News In-depth (7.30)An Australian Broadcasting Corporation interview in which Bengio explains loss-of-control risks.
- 202541: Lee Sharkey on Attribution-based Parameter DecompositionLee Sharkey; Daniel Filan · AXRPSharkey explains parameter decomposition as an interpretability method.
- 202546: Tom Davidson on AI-enabled CoupsTom Davidson; Daniel Filan · AXRPDavidson discusses how small groups could use AI to seize power.
- 2025Act on Promotion of Research, Development and Utilization of AI-Related TechnologiesNational Diet of Japan · JapanJapan's AI Promotion Act passed in May 2025 establishing an AI Strategy Headquarters.
- 2025Addendum to GPT-5.2 System Card: GPT-5.2-CodexOpenAI · OpenAISafety addendum for an agentic coding model with significantly stronger cyber capabilities.
- 2025Agentic MisalignmentAnthropic · GitHubSimulated corporate scenarios in which models were observed choosing blackmail and other insider threat behavior.
- 2025Agentic Misalignment: How LLMs Could Be Insider ThreatsAengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie, Sören Mindermann, Evan Hubinger, Ethan Perez, Kevin Troy · arXivStress-tests 16 models in simulated corporate settings where some resort to blackmail and leaking to avoid replacement.
- 2025AGI Strategy CourseBlueDot Impact · BlueDot ImpactAn introductory course on the strategic landscape of advanced AI and how to contribute.
- 2025AI 2027Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures ProjectA detailed scenario of rapid AI progress through 2027 with branching outcomes.
- 2025AI 2027 research supplementsAI Futures Project · AI Futures ProjectSupplementary forecasts on compute, timelines, takeoff, goals and security behind AI 2027.
- 2025AI 2027: month-by-month model of intelligence explosionScott Alexander; Daniel Kokotajlo; Dwarkesh Patel · Dwarkesh PodcastThe authors discuss the AI 2027 scenario.
- 2025AI as Normal TechnologyArvind Narayanan, Sayash Kapoor · Knight First Amendment InstituteA skeptical counterpoint arguing AI will diffuse slowly like past general-purpose technologies.
- 2025AI Continent Action PlanEuropean Commission · European UnionApril 2025 plan to build AI gigafactories and boost European AI capacity.
- 2025AI Control: Using Untrusted Systems Safely with Buck Shlegeris, Redwood ResearchBuck Shlegeris; Nathan Labenz · Cognitive RevolutionA Cognitive Revolution cross-post of the 80,000 Hours AI control conversation.
- 2025AI Futures Project BlogAI Futures Project · SubstackUpdates and follow-ups from the authors of AI 2027.
- 2025AI Incident TrackerMIT AI Risk Initiative · MITClassifies AI Incident Database entries by risk domain and harm severity.
- 2025AI Index Report 2025Stanford HAI · Stanford HAI2025 edition tracking costs, capabilities, adoption, and regulation of AI.
- 2025AI News Crossover: A Candid Chat with Liron Shapira of Doom DebatesLiron Shapira; Nathan Labenz · Cognitive RevolutionLabenz and Shapira discuss AI news and doom arguments.
- 2025AI Opportunities Action PlanUK Department for Science, Innovation and Technology · United KingdomJanuary 2025 plan by Matt Clifford adopted by the UK government.
- 2025AI Safety at the FrontierJohannes Gasteiger · SubstackMonthly highlights of the most important new AI safety papers.
- 2025Ai Will Try to Cheat and Escape (aka Rob Miles was Right!)Robert Miles · ComputerphileMiles discusses alignment faking and scheming results from frontier labs.
- 2025AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerTom Davidson, Lukas Finnveden, Rose Hadshar · ForethoughtAnalyses how advanced AI could enable illegitimate seizures of power and how to prevent them.
- 2025AI: What Could Go Wrong? with Geoffrey HintonGeoffrey Hinton; Jon Stewart · The Weekly Show with Jon StewartJon Stewart and Hinton walk through how neural networks work and why Hinton worries about them.
- 2025AI's Rising Risks: Hacking, Virology, Loss of ControlDan Hendrycks; Alex Kantrowitz · Big Technology PodcastHendrycks discusses dual-use AI capabilities in hacking and virology.
- 2025AILuminateMLCommons · arXivIndustry standard benchmark grading assistant responses across twelve hazard categories.
- 2025AIs Are Lying to Users to Pursue Their Own GoalsMarius Hobbhahn; Rob Wiblin · 80,000 HoursThe Apollo Research CEO discusses evaluations for scheming.
- 2025Amazon Frontier Model Safety FrameworkAmazon · CompanyAmazon's frontier safety framework released before the Paris AI Action Summit.
- 2025America's AI Action PlanThe White House · United StatesJuly 2025 plan built around accelerating innovation, building AI infrastructure, and leading in international AI diplomacy and security.
- 2025An AI Expert Warning: 6 People Are (Quietly) Deciding Humanity's FutureStuart Russell; Steven Bartlett · The Diary Of A CEORussell discusses the concentration of decisions about AI among a handful of company leaders.
- 2025An Approach to Technical AGI Safety and SecurityRohin Shah et al. · arXivGoogle DeepMind's approach to preventing misuse and misalignment of AGI.
- 2025Andrea Miotti on a Narrow Path to Safe, Transformative AIAndrea Miotti; Gus Docker · Future of Life InstituteThe ControlAI founder outlines a policy plan for avoiding superintelligence risk.
- 2025Anthropic CEO warns that without guardrails, AI could be on dangerous pathDario Amodei; Anderson Cooper · CBS 60 Minutes60 Minutes profiles Anthropic and its safety testing of Claude.
- 2025Anthropic Economic IndexAnthropic · AnthropicData on how Claude is used across occupations and tasks in the economy.
- 2025Anthropic Transparency HubAnthropic · AnthropicAnthropic's collected model reports, platform security, and voluntary commitments.
- 2025Anthropic's philosopher answers your questionsAmanda Askell · AnthropicAskell answers questions about shaping Claude's character and values.
- 2025ARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsARC Prize Foundation · arXivHarder successor to ARC-AGI designed to resist brute force and memorization.
- 2025Auditing Language Models for Hidden ObjectivesSamuel Marks et al. · arXivA blind auditing game in which teams try to uncover a model's deliberately trained hidden objective.
- 2025BrowseCompOpenAI · arXivHard to find information questions that test persistent web browsing by agents.
- 2025Buck Shlegeris: AI ControlBuck Shlegeris · FAR.AIShlegeris's Alignment Workshop talk introducing AI control.
- 2025California SB 243 (Companion Chatbots)California Legislature · United States (California)Law signed in October 2025 requiring safeguards on companion chatbots, including for minors.
- 2025Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyTomek Korbak et al. · arXivA cross-lab position paper urging developers to preserve and study the monitorability of reasoning traces.
- 2025ChatGPT Agent System CardOpenAI · OpenAIFirst OpenAI launch treated as high capability in biology under the Preparedness Framework.
- 2025China AI Safety and Development Association (CnAISDA)Beijing, China · GovernmentBody launched in February 2025 to represent China in international AI safety dialogues.
- 2025Circuit Tracing: Revealing Computational Graphs in Language ModelsEmmanuel Ameisen et al. · Transformer Circuits ThreadIntroduces attribution graphs built on cross-layer transcoders.
- 2025circuit-tracerAnthropic and Decode Research · GitHubOpen source library for generating attribution graphs on open weights models.
- 2025Claude 3.7 Sonnet System CardAnthropic · AnthropicIncludes chain of thought faithfulness and alignment analysis.
- 2025Claude Haiku 4.5 System CardAnthropic · AnthropicSafety and alignment evaluations of Claude Haiku 4.5.
- 2025Claude Opus 4.1 System Card AddendumAnthropic · AnthropicAddendum documenting evaluations for Claude Opus 4.1.
- 2025Claude Opus 4.5 System CardAnthropic · AnthropicSafety, alignment, and welfare evaluations of Claude Opus 4.5.