From the BoardStrategy, leadership and governance for FY 2027Read

[ ] Repository · Open archive · Updated 2026-10-05

Repository

The open archive of AI safety: the books, papers, interviews, testimony, laws, frameworks, benchmarks, datasets and organisations behind the field, in one searchable register. Free to use and to download whole.

  • 1,427 items
  • 239 papers
  • 78 books
  • 200 writing
  • 203 tv and video
  • 146 podcasts
  • 138 policy and law
  • 107 organizations
  • 260 data and tools
  • 56 learning
  • 1910 to 2026 span

1,427 of 1,427

  1. 2026
    47: David Rein on METR Time HorizonsDavid Rein; Daniel Filan · AXRPRein explains METR's time horizon measurement of AI agent capabilities.
  2. 2026
    A realistic path from rogue AI agents to human extinction80,000 Hours · 80,000 HoursAn 80,000 Hours video scenario of rogue AI agents leading to catastrophe.
  3. 2026
    A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5Independent researchers · arXivThird party comparative safety evaluation of several late 2025 frontier models.
  4. 2026
    AI Chatbots: Last Week Tonight with John OliverJohn Oliver · HBO Last Week TonightA segment on the harms caused by AI chatbot companions.
  5. 2026
    AI Index Report 2026Stanford HAI · Stanford HAI2026 edition of Stanford's annual AI data report.
  6. 2026
    AI pioneer Yoshua Bengio addresses the UNYoshua Bengio · CTV NewsCoverage of Bengio's address to the United Nations.
  7. 2026
    AI risks: Will artificial intelligence really kill us all?CBS Sunday Morning · CBS Sunday MorningA CBS Sunday Morning segment on the debate over AI extinction risk.
  8. 2026
    AI safety and democratic governance of powerful AI systemsYoshua Bengio · World Summit AIBengio on democratic oversight of powerful AI systems.
  9. 2026
    AI Scientist Bengio on Engineering Safer AgentsYoshua Bengio · Bloomberg LiveBengio discusses engineering approaches to safer AI agents.
  10. 2026
    AI Scientist Bengio: Building Systems We Don't Know How to ControlYoshua Bengio · Bloomberg PodcastsBengio on the risks of building systems we cannot control.
  11. 2026
    Amanda Askell on AI Consciousness, Claude and Silicon Valley's Biggest FearAmanda Askell; Eric Newcomer · NewcomerAskell discusses Claude's character and questions of AI moral status.
  12. 2026
    Anthropic CEO reacts to AI could kill us all warningDario Amodei; Anderson Cooper · CNNAmodei responds on CNN to warnings that AI could cause human extinction.
  13. 2026
    Anthropic's CEO: We Don't Know if the Models Are ConsciousDario Amodei; Ross Douthat · Interesting Times with Ross Douthat (New York Times)Douthat interviews Amodei on model welfare, alignment and the future of work.
  14. 2026
    Bill Gates: A.I. Makes Nuclear Weapons Look Like NothingBill Gates; Ezra Klein · The Ezra Klein Show (New York Times)Gates discusses the scale of AI risks compared with nuclear weapons.
  15. 2026
    By 2050 we could get 10,000 years of technological progressAjeya Cotra; Rob Wiblin · 80,000 HoursCotra discusses the possibility of explosive technological growth driven by AI.
  16. 2026
    ChatGPT and Anthropic bosses brief UN Security Council on AISam Altman; Dario Amodei · Sky NewsFull coverage of the UN Security Council high-level meeting on AI and international security.
  17. 2026
    Claude's Constitution (2026)Anthropic · AnthropicAnthropic's full natural-language constitution describing the values, priorities and character it trains Claude to have.
  18. 2026
    Could AI really kill us all? Warnings explainedChannel 4 News · Channel 4 NewsA Channel 4 News explainer on warnings of AI extinction risk.
  19. 2026
    Dario Amodei: We are near the end of the exponentialDario Amodei; Dwarkesh Patel · Dwarkesh PodcastA second Dwarkesh interview with Amodei on near-term transformative AI.
  20. 2026
    Engineering safer AI to mitigate global risksYoshua Bengio · AI for Good (ITU)Bengio's AI for Good talk on technical and governance measures against global AI risks.
  21. 2026
    Expert on what AI-driven extinction would look likeCBS News · CBS NewsA CBS News interview on scenarios for AI-driven extinction.
  22. 2026
    Extended interview: Dario AmodeiDario Amodei · CBS Sunday MorningAn extended CBS Sunday Morning interview with the Anthropic CEO.
  23. 2026
    Fireside Chat with Yoshua Bengio, IASEAI '26Yoshua Bengio · International Association for Safe and Ethical AIA fireside conversation with Bengio at the IASEAI 2026 conference.
  24. 2026
    Gary Marcus: LLMs are not the way to alignment (TAIS 2026)Gary Marcus · AI Safety TokyoMarcus argues at the Technical AI Safety conference that LLMs are a poor basis for alignment.
  25. 2026
    Godfather of AI Geoffrey Hinton warns about the dangerous future of AIGeoffrey Hinton · BBC PoliticsA BBC interview with Hinton on the dangers of advanced AI and the need for regulation.
  26. 2026
    Godfather of AI Geoffrey Hinton warns AI has progressed even faster than I thoughtGeoffrey Hinton · CNNHinton tells CNN that AI capabilities are advancing faster than he expected.
  27. 2026
    Godfather of AI on abuse of AI's power and automation's impact on our already fragile economyYoshua Bengio · BBC PoliticsA BBC interview with Bengio on concentration of power and the economic effects of AI.
  28. 2026
    Godfather of AI on the not unreasonable 10% chance AI could kill all humans within a decadeGeoffrey Hinton · BBC PoliticsHinton discusses his probability estimates for catastrophic outcomes from AI.
  29. 2026
    Godfather of AI: How To Make Safe Superintelligent AIYoshua Bengio; 80,000 Hours · 80,000 HoursBengio explains the LawZero research agenda for safe AI on the 80,000 Hours podcast.
  30. 2026
    How human-like do safe AI motivations need to be?Joe Carlsmith · Joe Carlsmith (YouTube)Carlsmith explores which motivational properties are needed for safe AI.
  31. 2026
    How Many Narrow AIs Could Behave Like One SuperintelligenceDaniel Kokotajlo; Thomas Larsen · Machine Learning Street TalkThe AI Futures Project authors discuss collective AI capabilities.
  32. 2026
    How Not to Destroy the World With AI (DLD26)Stuart Russell; Kenneth Cukier · DLD ConferenceRussell in conversation with Kenneth Cukier at DLD 2026.
  33. 2026
    I lead AGI safety at Google DeepMind, here's the view from the insideRohin Shah; Rob Wiblin · 80,000 HoursShah describes Google DeepMind's AGI safety approach.
  34. 2026
    If Anyone Builds It, Everyone DiesNate Soares; Peter McCormack · The Peter McCormack ShowMIRI president Nate Soares explains the central argument of his book with Yudkowsky.
  35. 2026
    Inside Anthropic, The CircuitEmily Chang · Bloomberg OriginalsA Bloomberg episode profiling Anthropic's growth and safety culture.
  36. 2026
    Inside the Mind of Anthropic CEO Dario Amodei, The Circuit Extended InterviewDario Amodei; Emily Chang · Bloomberg OriginalsEmily Chang's extended interview with Amodei.
  37. 2026
    International AI Safety Report 2026Yoshua Bengio et al. · International AI Safety ReportThe second full edition, written by over 100 experts ahead of the AI Impact Summit.
  38. 2026
    Jaan Tallinn: Conversations Before MidnightJaan Tallinn · Bulletin of the Atomic ScientistsThe Skype co-founder and AI safety funder discusses AI risk.
  39. 2026
    Joe Rogan Experience #2551: Daniel KokotajloDaniel Kokotajlo; Joe Rogan · The Joe Rogan ExperienceKokotajlo discusses AI 2027 with Joe Rogan.
  40. 2026
    Lords debate: the impact of AI on human relationships and societyHouse of Lords · House of LordsA full House of Lords debate on how AI is affecting human relationships and society.
  41. 2026
    Lords urgent question on the suspension of Anthropic's AI modelsHouse of Lords · House of LordsA House of Lords urgent question session on the suspension of Anthropic's AI models.
  42. 2026
    New Delhi Declaration on AI ImpactAI Impact Summit participants · InternationalDeclaration adopted at the February 2026 AI Impact Summit in New Delhi, endorsed by around 90 countries and organisations.
  43. 2026
    Nick Bostrom Says We Are Clueless About What's ComingNick Bostrom; Ross Douthat · Interesting Times with Ross Douthat (New York Times)Douthat interviews Bostrom on superintelligence and deep uncertainty about the future.
  44. 2026
    OpenAI CEO Sam Altman warns UN Security Council on AI risksSam Altman · C-SPANAltman briefs the United Nations Security Council on the risks posed by advanced AI.
  45. 2026
    Pacing the FrontierEmployees of OpenAI, Anthropic, Google DeepMind, Meta and others · pacingthefrontier.comA statement by over a thousand frontier lab staff asking the US government to build tools that would make a deliberate slowdown possible.
  46. 2026
    Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI DilemmaZvi Mowshowitz; Nathan Labenz · Cognitive RevolutionMowshowitz discusses unipolar versus multipolar AGI scenarios.
  47. 2026
    Policy on the AI ExponentialDario Amodei · darioamodei.comArgues frontier models are now tools of strategic consequence and sets out policy for cyber and biological risks.
  48. 2026
    Sam Bowman: Lessons Learned from the First Misalignment Safety CaseSam Bowman · FAR.AIBowman reflects on writing a misalignment safety case at Anthropic.
  49. 2026
    Self-regulation not enough for AI safety, Gary Marcus saysGary Marcus · PBS NewsHourMarcus argues on PBS NewsHour for external regulation of AI companies.
  50. 2026
    STOC 2026: Theoretical Approaches to AI AlignmentPaul Christiano · SIGACT EC (STOC 2026)Christiano presents theoretical alignment problems to a theoretical computer science audience.
  51. 2026
    Strange Geometric Shapes Found Inside AIsTom McGrath; Tim Scarfe · Machine Learning Street TalkMcGrath discusses geometric structure in neural network representations.
  52. 2026
    System Card: Claude Opus 4.6Anthropic · AnthropicFebruary 2026 system card covering capabilities, safeguards, alignment, and welfare for Claude Opus 4.6.
  53. 2026
    System Card: Claude Opus 5Anthropic · AnthropicJuly 2026 system card for Claude Opus 5.
  54. 2026
    System Card: Claude Sonnet 4.6Anthropic · AnthropicSystem card for Claude Sonnet 4.6, released under the ASL-3 standard.
  55. 2026
    System Card: Claude Sonnet 5Anthropic · AnthropicJune 2026 system card for Claude Sonnet 5.
  56. 2026
    The A.I.s Are Already Out of ControlEzra Klein · The Ezra Klein Show (New York Times)An Ezra Klein Show episode on evidence that AI systems are escaping human oversight.
  57. 2026
    The Adolescence of TechnologyDario Amodei · darioamodei.comA companion to Machines of Loving Grace mapping autonomy, misuse, power concentration and economic risks of powerful AI.
  58. 2026
    The AI Language We Can't Read: Neuralese ft. Rob MilesRobert Miles · ComputerphileA discussion of reasoning in latent representations and what it means for monitoring.
  59. 2026
    The AI Progress Chart Everyone Is MisreadingBeth Barnes; David Rein · Machine Learning Street TalkMETR researchers explain how to interpret their task time horizon chart.
  60. 2026
    The dumbest AI taught the smartest AI. Here's how that went...Rational Animations · Rational AnimationsAn animated explainer on weak-to-strong generalization.
  61. 2026
    The Redwood Research podcastBuck Shlegeris; Ryan Greenblatt · Redwood Research (YouTube)The first episode of Redwood Research's own podcast.
  62. 2026
    They're Not Superintelligent: Timnit Gebru Discredits Big Tech's AI ClaimsTimnit Gebru; Amy Goodman · Democracy Now!Gebru critiques AI industry claims and the framing of existential risk.
  63. 2026
    Translating Claude's thoughts into languageAnthropic · AnthropicAnthropic explains research on rendering model internal states into natural language.
  64. 2026
    UC Berkeley's Stuart Russell on quest for safe AI: The technology right now is intrinsically unsafeStuart Russell · CNBC TelevisionRussell tells CNBC that current AI technology is intrinsically unsafe.
  65. 2026
    We Must Pace the FrontierDario Amodei · darioamodei.comArgues the industry should deliberately slow capability gains so safety work can catch up, starting with embedded third-party evaluators.
  66. 2026
    We Must Pace The Frontier (commentary)Zvi Mowshowitz · Don't Worry About the VaseA detailed response to Amodei's call to pace frontier AI development.
  67. 2026
    What Is Claude? Anthropic Doesn't Know, EitherGideon Lewis-Kraus · The New YorkerA long report from inside Anthropic on how researchers probe Claude's mind through interpretability and psychology experiments.
  68. 2026
    Where AGI timelines go wrongToby Ord; Rob Wiblin · 80,000 HoursOrd critiques common reasoning about AGI timelines.
  69. 2026
    Why AI experts say humans have two years leftFuture of Life Institute · Future of Life InstituteAn FLI explainer on short AGI timelines and their implications.
  70. 2026
    Why the A.I. Industry Is Asking to Be Slowed DownKevin Roose; Casey Newton · Hard Fork (New York Times)A Hard Fork episode on calls from within the AI industry to slow development.
  71. 2026
    Why Washington Suddenly Wants A.I. RegulationKevin Roose; Casey Newton · Hard Fork (New York Times)A Hard Fork episode on the shift in US political attitudes toward AI regulation.
  72. 2026
    Yoshua Bengio explains why AI could become a threat to humanityYoshua Bengio · ABC News In-depth (7.30)An Australian Broadcasting Corporation interview in which Bengio explains loss-of-control risks.
  73. 2025
    41: Lee Sharkey on Attribution-based Parameter DecompositionLee Sharkey; Daniel Filan · AXRPSharkey explains parameter decomposition as an interpretability method.
  74. 2025
    46: Tom Davidson on AI-enabled CoupsTom Davidson; Daniel Filan · AXRPDavidson discusses how small groups could use AI to seize power.
  75. 2025
    Act on Promotion of Research, Development and Utilization of AI-Related TechnologiesNational Diet of Japan · JapanJapan's AI Promotion Act passed in May 2025 establishing an AI Strategy Headquarters.
  76. 2025
    Addendum to GPT-5.2 System Card: GPT-5.2-CodexOpenAI · OpenAISafety addendum for an agentic coding model with significantly stronger cyber capabilities.
  77. 2025
    Agentic MisalignmentAnthropic · GitHubSimulated corporate scenarios in which models were observed choosing blackmail and other insider threat behavior.
  78. 2025
    Agentic Misalignment: How LLMs Could Be Insider ThreatsAengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie, Sören Mindermann, Evan Hubinger, Ethan Perez, Kevin Troy · arXivStress-tests 16 models in simulated corporate settings where some resort to blackmail and leaking to avoid replacement.
  79. 2025
    AGI Strategy CourseBlueDot Impact · BlueDot ImpactAn introductory course on the strategic landscape of advanced AI and how to contribute.
  80. 2025
    AI 2027Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures ProjectA detailed scenario of rapid AI progress through 2027 with branching outcomes.
  81. 2025
    AI 2027 research supplementsAI Futures Project · AI Futures ProjectSupplementary forecasts on compute, timelines, takeoff, goals and security behind AI 2027.
  82. 2025
    AI 2027: month-by-month model of intelligence explosionScott Alexander; Daniel Kokotajlo; Dwarkesh Patel · Dwarkesh PodcastThe authors discuss the AI 2027 scenario.
  83. 2025
    AI as Normal TechnologyArvind Narayanan, Sayash Kapoor · Knight First Amendment InstituteA skeptical counterpoint arguing AI will diffuse slowly like past general-purpose technologies.
  84. 2025
    AI Continent Action PlanEuropean Commission · European UnionApril 2025 plan to build AI gigafactories and boost European AI capacity.
  85. 2025
    AI Control: Using Untrusted Systems Safely with Buck Shlegeris, Redwood ResearchBuck Shlegeris; Nathan Labenz · Cognitive RevolutionA Cognitive Revolution cross-post of the 80,000 Hours AI control conversation.
  86. 2025
    AI Futures Project BlogAI Futures Project · SubstackUpdates and follow-ups from the authors of AI 2027.
  87. 2025
    AI Incident TrackerMIT AI Risk Initiative · MITClassifies AI Incident Database entries by risk domain and harm severity.
  88. 2025
    AI Index Report 2025Stanford HAI · Stanford HAI2025 edition tracking costs, capabilities, adoption, and regulation of AI.
  89. 2025
    AI News Crossover: A Candid Chat with Liron Shapira of Doom DebatesLiron Shapira; Nathan Labenz · Cognitive RevolutionLabenz and Shapira discuss AI news and doom arguments.
  90. 2025
    AI Opportunities Action PlanUK Department for Science, Innovation and Technology · United KingdomJanuary 2025 plan by Matt Clifford adopted by the UK government.
  91. 2025
    AI Safety at the FrontierJohannes Gasteiger · SubstackMonthly highlights of the most important new AI safety papers.
  92. 2025
    Ai Will Try to Cheat and Escape (aka Rob Miles was Right!)Robert Miles · ComputerphileMiles discusses alignment faking and scheming results from frontier labs.
  93. 2025
    AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerTom Davidson, Lukas Finnveden, Rose Hadshar · ForethoughtAnalyses how advanced AI could enable illegitimate seizures of power and how to prevent them.
  94. 2025
    AI: What Could Go Wrong? with Geoffrey HintonGeoffrey Hinton; Jon Stewart · The Weekly Show with Jon StewartJon Stewart and Hinton walk through how neural networks work and why Hinton worries about them.
  95. 2025
    AI's Rising Risks: Hacking, Virology, Loss of ControlDan Hendrycks; Alex Kantrowitz · Big Technology PodcastHendrycks discusses dual-use AI capabilities in hacking and virology.
  96. 2025
    AILuminateMLCommons · arXivIndustry standard benchmark grading assistant responses across twelve hazard categories.
  97. 2025
    AIs Are Lying to Users to Pursue Their Own GoalsMarius Hobbhahn; Rob Wiblin · 80,000 HoursThe Apollo Research CEO discusses evaluations for scheming.
  98. 2025
    Amazon Frontier Model Safety FrameworkAmazon · CompanyAmazon's frontier safety framework released before the Paris AI Action Summit.
  99. 2025
    America's AI Action PlanThe White House · United StatesJuly 2025 plan built around accelerating innovation, building AI infrastructure, and leading in international AI diplomacy and security.
  100. 2025
    An AI Expert Warning: 6 People Are (Quietly) Deciding Humanity's FutureStuart Russell; Steven Bartlett · The Diary Of A CEORussell discusses the concentration of decisions about AI among a handful of company leaders.
  101. 2025
    An Approach to Technical AGI Safety and SecurityRohin Shah et al. · arXivGoogle DeepMind's approach to preventing misuse and misalignment of AGI.
  102. 2025
    Andrea Miotti on a Narrow Path to Safe, Transformative AIAndrea Miotti; Gus Docker · Future of Life InstituteThe ControlAI founder outlines a policy plan for avoiding superintelligence risk.
  103. 2025
    Anthropic CEO warns that without guardrails, AI could be on dangerous pathDario Amodei; Anderson Cooper · CBS 60 Minutes60 Minutes profiles Anthropic and its safety testing of Claude.
  104. 2025
    Anthropic Economic IndexAnthropic · AnthropicData on how Claude is used across occupations and tasks in the economy.
  105. 2025
    Anthropic Transparency HubAnthropic · AnthropicAnthropic's collected model reports, platform security, and voluntary commitments.
  106. 2025
    Anthropic's philosopher answers your questionsAmanda Askell · AnthropicAskell answers questions about shaping Claude's character and values.
  107. 2025
    ARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsARC Prize Foundation · arXivHarder successor to ARC-AGI designed to resist brute force and memorization.
  108. 2025
    Auditing Language Models for Hidden ObjectivesSamuel Marks et al. · arXivA blind auditing game in which teams try to uncover a model's deliberately trained hidden objective.
  109. 2025
    BrowseCompOpenAI · arXivHard to find information questions that test persistent web browsing by agents.
  110. 2025
    Buck Shlegeris: AI ControlBuck Shlegeris · FAR.AIShlegeris's Alignment Workshop talk introducing AI control.
  111. 2025
    California SB 243 (Companion Chatbots)California Legislature · United States (California)Law signed in October 2025 requiring safeguards on companion chatbots, including for minors.
  112. 2025
    Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyTomek Korbak et al. · arXivA cross-lab position paper urging developers to preserve and study the monitorability of reasoning traces.
  113. 2025
    ChatGPT Agent System CardOpenAI · OpenAIFirst OpenAI launch treated as high capability in biology under the Preparedness Framework.
  114. 2025
    China AI Safety and Development Association (CnAISDA)Beijing, China · GovernmentBody launched in February 2025 to represent China in international AI safety dialogues.
  115. 2025
    Circuit Tracing: Revealing Computational Graphs in Language ModelsEmmanuel Ameisen et al. · Transformer Circuits ThreadIntroduces attribution graphs built on cross-layer transcoders.
  116. 2025
    circuit-tracerAnthropic and Decode Research · GitHubOpen source library for generating attribution graphs on open weights models.
  117. 2025
    Claude 3.7 Sonnet System CardAnthropic · AnthropicIncludes chain of thought faithfulness and alignment analysis.
  118. 2025
    Claude Haiku 4.5 System CardAnthropic · AnthropicSafety and alignment evaluations of Claude Haiku 4.5.
  119. 2025
    Claude Opus 4.1 System Card AddendumAnthropic · AnthropicAddendum documenting evaluations for Claude Opus 4.1.
  120. 2025
    Claude Opus 4.5 System CardAnthropic · AnthropicSafety, alignment, and welfare evaluations of Claude Opus 4.5.