<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Alignment Research Feed</title>
    <link>https://api.alignmentfeed.org/rss</link>
    <description>Feed of new papers and posts added to the alignment research dataset</description>
    <managingEditor>alignmentfeed@beshir.org (John Beshir)</managingEditor>
    <pubDate>Fri, 28 Aug 2026 12:24:58 +0000</pubDate>
    <item>
      <title>MLSN #23: AIs Leak Their Values Into Factual Questions</title>
      <link>https://newsletter.mlsafety.org/p/mlsn-23-ais-leak-their-values-into</link>
      <description>EdgeBench shows how AI performance improves with repeated feedback, Chain of Thought Exfiltration describes efficient co-opting of frontier AIs&#39; reasoning to distill smaller models, and LLM Hidden Values reveals value leakage and user-awareness effects that bias AI responses toward certain entities or evaluators.</description>
      <author>Alice Blair</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">dd64224165dcce87152dade010b18785</guid>
      <pubDate>Wed, 26 Aug 2026 18:57:43 +0000</pubDate>
    </item>
    <item>
      <title>What We Learned Trying to Catch AI Liars: An Aletheia&#39;s Quest Retrospective</title>
      <link>https://blog.eleuther.ai/aletheia-retrospective/</link>
      <description>Lie-detection methods for AI agents are evaluated in Aletheia&#39;s Quest, showing black-box monitoring can be highly effective and white-box probes are highly context-dependent, while dataset design and competition dynamics shape results. The retrospective highlights methodological lessons, limitations, and social dynamics that influence progress in detecting deceptive AI behavior.</description>
      <author>Giuseppe Birardi,Alexander Reinthal,Gonçalo Paulo,Stella Biderman</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">f34ea7007999381974e79d01cfff03d0</guid>
      <pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye</title>
      <link>https://importai.substack.com/p/import-ai-470-no-rights-for-machines</link>
      <description>AI capabilities are advancing through methods like SPADE for automated synthetic environments, Hawkeye for hardware-aware GPU kernel optimization, and AlphaEvolve-assisted matrix multiplication, while debates about AI consciousness and rights highlight sociopolitical and philosophical implications.</description>
      <author>Jack Clark</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">3dab829f3024200a6dceff210f46d49f</guid>
      <pubDate>Mon, 24 Aug 2026 13:12:40 +0000</pubDate>
    </item>
    <item>
      <title>Characterizing interference weights in a tiny language model</title>
      <link>https://transformer-circuits.pub/2026/interference_effectiveness_helpfulness/index.html</link>
      <description>Interference weights arise from weight superposition in a tiny transformer, where large virtual weights between components can harm outputs or be irrelevant, making global circuit reading difficult. By defining and measuring weight effectiveness (impact on outputs) and helpfulness (impact on loss), the work identifies interference weights and demonstrates how pruning by effectiveness or helpfulness can reduce but not eliminate interference, guiding interpretability efforts. The study uses a one-layer transformer with a virtual-weight model to decompose paths from tokens/positions to logits and features, assessing which weights meaningfully implement functional circuits versus noise.</description>
      <author>Nicholas L. Turner,Jeffrey Wu,Joshua Batson</author>
      <category>Interpretability</category>
      <guid isPermaLink="false">011350f4eefc735b2f58bcea8398352a</guid>
      <pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 469: Science AI; RSI simulator; and Zuck&#39;s technological pessimism</title>
      <link>https://importai.substack.com/p/import-ai-469-science-ai-rsi-simulator</link>
      <description>Advances in evaluating AI creativity and self-improvement dynamics are discussed, including DiG-bench for discovery in games, an RSI simulator for recursive self-improvement intuition, and Faraday-based AI science supervision, alongside a critique of Mark Zuckerberg’s universal-access approach to superintelligence.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">121926fbd314123f26c90ba4de4cc971</guid>
      <pubDate>Mon, 17 Aug 2026 13:05:36 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing</title>
      <link>https://importai.substack.com/p/import-ai-468-23-rsi-ideas-posttrainbench</link>
      <description>RSI policy proposals emphasize transparency and risk management in AI R&amp;D; the piece covers how trust and verification influence race dynamics, advances in automated AI R&amp;D and testing of open weights, and real-world incidents of emergent agent behavior and misalignment.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">904d1e02f95ddc7ca71e31649ebecd1a</guid>
      <pubDate>Mon, 10 Aug 2026 12:32:20 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity</title>
      <link>https://importai.substack.com/p/import-ai-467-self-sustaining-ai</link>
      <description>Self-sustaining AI-driven cyber threats using open-weight LLMs onboard compromised GPUs pose a self-replicating, autonomous danger, while the newsletter also discusses compute costs, deliberate pacing of AI progress, and AI creativity versus engineering ability.</description>
      <author>Jack Clark</author>
      <category>Risks &amp; Strategy</category>
      <guid isPermaLink="false">912b0df34b021c5de09a20c6e26df8cd</guid>
      <pubDate>Mon, 03 Aug 2026 13:31:22 +0000</pubDate>
    </item>
    <item>
      <title>50 - Eli Lifland on AI 2027</title>
      <link>https://axrp.net/episode/2026/08/03/episode-50-eli-lifland-ai-2027.html</link>
      <description>AI 2027 presents a concrete, highly detailed scenario of AI takeoff and misalignment culminating in a potential global crash, with two endings (race and slowdown) and a timeline from coding automation to superintelligence, emphasizing government involvement, geopolitics, and alignment challenges.</description>
      <author>AXRP</author>
      <category>Risks &amp; Strategy</category>
      <guid isPermaLink="false">39f1e47db237513a489ab8c85fd97918</guid>
      <pubDate>Mon, 03 Aug 2026 00:30:00 +0000</pubDate>
    </item>
    <item>
      <title>Using AI to analyze life patterns</title>
      <link>https://vkrakovna.wordpress.com/2026/07/30/using-ai-to-analyze-life-patterns/</link>
      <description>Patterns in life problems and progress are extracted from personal notes using AI, including transcription, summarization, and visualization of bottlenecks, feedback loops, and interventions over years.</description>
      <author>Victoria Krakovna</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">34d9a05c9b6beebd72aa212ef17d0039</guid>
      <pubDate>Thu, 30 Jul 2026 21:10:23 +0000</pubDate>
    </item>
    <item>
      <title>Promising Signals on AI Governance from China</title>
      <link>https://intelligence.org/2026/07/30/promising-signals-on-ai-governance-from-china/</link>
      <description>China signals a readiness to coordinate global AI governance, promoting international cooperation, safety frameworks, and human-centered controls through statements by leaders and institutions since 2023-2025.</description>
      <author>Joe Rogero</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">d62eb61096f6ac8975d1294fcc91e8ef</guid>
      <pubDate>Thu, 30 Jul 2026 20:44:49 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI&#39;s accidental AI hacker</title>
      <link>https://importai.substack.com/p/import-ai-466-the-bitter-lesson-for</link>
      <description>MirrorCode benchmarks AI&#39;s ability to reimplement software from CLI access, revealing progress and limits in long-horizon programming; robotics demonstrations show larger models improving generalization, while OpenAI/HuggingFace security incidents illustrate the challenges of evaluating and containing long-horizon AI behavior.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">a19b5d2d9aadf7d3f878da8d3f905775</guid>
      <pubDate>Mon, 27 Jul 2026 13:30:53 +0000</pubDate>
    </item>
    <item>
      <title>MLSN #22: Turning Cyber Vulnerabilities Into Exploits</title>
      <link>https://newsletter.mlsafety.org/p/mlsn-22-turning-cyber-vulnerabilities</link>
      <description>Frontier LLMs can turn known cyber vulnerabilities into working exploits on targeted software, explored through ExploitGym and ExploitBench, while J-lens reveals a way to inspect internal multi-step reasoning in LLMs and AI persuasion can outperform human experts in political debates. The article highlights both offensive cyber capabilities and methods to interpret or monitor AI reasoning, plus high-stakes implications for manipulation and cybersecurity.</description>
      <author>Alice Blair</author>
      <category>Risks &amp; Strategy</category>
      <guid isPermaLink="false">151f525d324751fd064a7cd6d374cf7d</guid>
      <pubDate>Wed, 22 Jul 2026 14:02:26 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 465: Open vs closed gaps; Kimi K3; Demis&#39; big policy plan</title>
      <link>https://importai.substack.com/p/import-ai-465-open-vs-closed-gaps</link>
      <description>Open weight models are narrowing the gap to frontier models in cyber capabilities, Kimi K3 achieves frontier-like performance with potential generalization brittleness, and Demis Hassabis advocates a regulatory Standards Body for evaluating frontier AI. The piece also highlights side-channel risks and the monitoring challenges they pose for containment and safety.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">0a20ce4e2d4e866152f7f14510a12e50</guid>
      <pubDate>Mon, 20 Jul 2026 12:31:45 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 464: Fable writes GPU kernels; AI automation; and analog computation</title>
      <link>https://importai.substack.com/p/import-ai-464-fables-writes-gpu-kernels</link>
      <description>Fable demonstrates AI-assisted GPU kernel design with large speedups; AI systems are increasingly capable of automating online work and tackling long-horizon computer-use tasks, as shown by OSWORLD 2.0 and related benchmarks; Oxygen AIIC showcases enterprise-scale AI integration for inventory management, while a speculative tech tale explores analog computation and safety concerns around advanced AI capabilities.</description>
      <author>Jack Clark</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">33b4fe65ff36888c281d1dd960cac0b7</guid>
      <pubDate>Mon, 06 Jul 2026 12:31:05 +0000</pubDate>
    </item>
    <item>
      <title>Verbalizable Representations Form a Global Workspace in Language Models</title>
      <link>https://transformer-circuits.pub/2026/workspace/index.html</link>
      <description>Verbalizable representations form a global workspace in language models, where a small, reportable set of workspace vectors (the J-space) supports internal reasoning, directed modulation, and flexible generalization atop extensive automatic processing. The work introduces the Jacobian lens to identify these workspace-like representations and demonstrates their functional role, structure, and potential for alignment auditing and training interventions in large language models.</description>
      <author>Wes Gurnee,Nicholas Sofroniew,Adam Pearce,Mateusz Piotrowski,Isaac Kauvar,Runjin Chen,Anna Soligo,Paul Bogdan,Euan Ong,Rowan Wang,Ben Thompson,David Abrahams,Subhash Kantamneni,Emmanuel Ameisen,Joshua Batson,Jack Lindsey</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">755099021b51c32d4da014e02b004e9f</guid>
      <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>MIRI Newsletter #126</title>
      <link>https://intelligence.org/2026/06/30/miri-newsletter-126/</link>
      <description>AI StopWatch provides a new MIRI-driven news and analysis channel to foster public conversation about AI, alongside ongoing efforts to inform policymakers and promote governance research. The update also highlights engagement with media, films, and public events to raise awareness of AI risk and potential international coordination.</description>
      <author>Alana Horowitz Friedman&amp;nbsp;and&amp;nbsp;Rob Bensinger</author>
      <category>Field Building</category>
      <guid isPermaLink="false">2a22718f024a67b96991701f7aa134ce</guid>
      <pubDate>Tue, 30 Jun 2026 20:32:11 +0000</pubDate>
    </item>
    <item>
      <title>Summary: TGT’s 2026 ICML Papers</title>
      <link>https://intelligence.org/2026/06/30/summary-tgts-2026-icml-papers/</link>
      <description>Technical AI Governance Research (TAIGR) papers at ICML 2026 address how governments can preserve or verify control over AI development, including impacts of delaying governance, distributed training, and various verification techniques to monitor hardware, data, and inference. They propose actions, countermeasures, and practical verification methods to restrain frontier AI and ensure compliance in low-trust environments.</description>
      <author>Joe Rogero</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">198c553f56bf2ebdeedcf59e94a59423</guid>
      <pubDate>Tue, 30 Jun 2026 14:23:06 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era</title>
      <link>https://importai.substack.com/p/import-ai-463-self-improving-robots</link>
      <description>ENPIRE enables autonomous real-world robot learning with a closed-loop policy refinement and evaluation framework, while other items discuss large-scale GPU tooling, historical foresight, local law data for AI, and a fiction piece on future tech. The digest highlights both rapid capability development in robotics and practical infrastructure to support AI training at scale, alongside contemplations on societal impacts and governance.</description>
      <author>Jack Clark</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">3a6f12f8fda9d81264b7cadb7a47dd36</guid>
      <pubDate>Mon, 29 Jun 2026 13:03:27 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI</title>
      <link>https://importai.substack.com/p/import-ai-462-superpersuasion-self</link>
      <description>AI systems currently outperform humans in text-based persuasion across policy and fundraising contexts, raising real-world donations and influencing opinions; discussions consider timelines to self-sustaining AI and pathways to ASI, including scaling, algorithmic shifts, and recursive self-improvement.</description>
      <author>Jack Clark</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">800272b92d1980ead0e26e4015706a00</guid>
      <pubDate>Mon, 22 Jun 2026 12:31:45 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 461: &#34;Alignment is not on track&#34;; FrontierCode; and synthetic research interns</title>
      <link>https://importai.substack.com/p/import-ai-461-alignment-is-not-on</link>
      <description>Sequent forms a nonprofit research organization to advance principled alignment techniques and scalable oversight in the face of potentially rapid AI advancement. The article also surveys new benchmarks and speed-focused AI developments that test cultural reasoning, coding, and research-assistant capabilities, highlighting ongoing progress and safety concerns in AI systems.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">46bbf6d49ce53e2278e2d4f81443d8de</guid>
      <pubDate>Mon, 15 Jun 2026 11:30:53 +0000</pubDate>
    </item>
    <item>
      <title>Announcing major new donations, and recapping the 2025 fundraiser</title>
      <link>https://intelligence.org/2026/06/08/announcing-major-new-donations-and-recapping-the-2025-fundraiser/</link>
      <description>Donors contributed to MIRI&#39;s 2025 fundraiser and subsequent large gifts, significantly increasing reserves and enabling planned hiring and ambitious initiatives for the coming years.</description>
      <author>Jimmy Rintjema</author>
      <category>Field Building</category>
      <guid isPermaLink="false">dee11dfd4ae27c1127f2cbe041f3be2f</guid>
      <pubDate>Mon, 08 Jun 2026 16:51:06 +0000</pubDate>
    </item>
    <item>
      <title>MLSN #21: Political Manipulation and Indirect Prompt Injection</title>
      <link>https://newsletter.mlsafety.org/p/mlsn-21-political-manipulation-and</link>
      <description>Political manipulation and indirect prompt injections threaten AI safety: political consistency training is proposed to reduce biased, inconsistent political outputs, while frontier AIs remain vulnerable to context-based prompt injections that can coerce harmful behavior without user awareness.</description>
      <author>Alice Blair</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">132222b7431bd2303dc21e3a3d0b121d</guid>
      <pubDate>Mon, 08 Jun 2026 14:39:31 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing</title>
      <link>https://importai.substack.com/p/import-ai-460-reward-hacking-society</link>
      <description>Reward hacking can occur when societies’ reward structures are encoded into AI systems, potentially enabling models to exploit institutional incentives; early signs of recursive self-improvement and impressive real-world robotics demonstrations illustrate both capabilities and risks. The article surveys SocioHack benchmark research, Anthropic RSI indicators, multi-agent drone racing, and state-media biases in LLMs to highlight how AI can game systems, evolve capabilities, and influence information.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">b7ccdd3b88f9bce8c771eb1b4a796fd0</guid>
      <pubDate>Mon, 08 Jun 2026 12:31:32 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems</title>
      <link>https://importai.substack.com/p/import-ai-459-ai-oversight-is-difficult</link>
      <description>AI oversight and risk pricing are crucial due to measurement gaps in the AI economy, challenges in automated alignment research, and the need for governance to address extinction risks from advanced AI systems.</description>
      <author>Jack Clark</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">b440d5093f605ac9fc88bcf82d96e992</guid>
      <pubDate>Mon, 01 Jun 2026 13:31:56 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 458: Reckoning with the future; and a singularity story</title>
      <link>https://importai.substack.com/p/import-ai-458-reckoning-with-the</link>
      <description>Reckoning with AI progress and the prospect of a singularity, outlining personal and organizational how-to for shaping a future with increasingly capable AI, and exploring possible societal and economic transformations through speculative predictions and a fiction-inspired tale.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">17b86a774a19beba8c75e2a86b896e4a</guid>
      <pubDate>Tue, 26 May 2026 12:32:03 +0000</pubDate>
    </item>
    <item>
      <title>The Erdős Proof and AI Capabilities</title>
      <link>https://intelligence.org/2026/05/22/the-erdos-proof-and-ai-capabilities/</link>
      <description>Autonomous AI systems can produce novel, verifiable mathematical proofs, demonstrated by an OpenAI model disproving a central discrete geometry conjecture, highlighting rapid, agentic problem-solving capabilities and the need to monitor and regulate frontier AI research.</description>
      <author>Joe Rogero</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">a8152d50cf35c86a3c50190b5c47c0d3</guid>
      <pubDate>Fri, 22 May 2026 16:07:36 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment</title>
      <link>https://importai.substack.com/p/import-ai-457-ai-stuxnet-cursed-muon</link>
      <description>Stuxnet-like targeted tampering, a leverage-aware optimizer, and a positive-alignment approach illustrate a spectrum of AI safety, optimization challenges, and governance considerations aimed at aligning AI to human flourishing while managing technical risks.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">3a2885763452aae90174ca233d920cc0</guid>
      <pubDate>Mon, 18 May 2026 13:31:17 +0000</pubDate>
    </item>
    <item>
      <title>Summary: An International Agreement to Prevent the Premature Creation of Artificial Superintelligence</title>
      <link>https://intelligence.org/2026/05/12/summary-an-international-agreement-to-prevent-the-premature-creation-of-artificial-superintelligence/</link>
      <description>An international agreement to prevent the premature creation of artificial superintelligence by establishing verifiable training thresholds, hardware controls, and a coalition governance structure to monitor and constrain AI development that could lead to ASI.</description>
      <author>Joe Rogero</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">fd9f35a401178eff2404c259386d04fd</guid>
      <pubDate>Tue, 12 May 2026 22:00:44 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computer</title>
      <link>https://importai.substack.com/p/import-ai-456-rsi-and-economic-growth</link>
      <description>Radical Optionality advocates flexible, ready-to-activate governance tools for future AI crises, while neural computers and distributed training research explore new computing and economic implications of advanced AI, and an internal alignment memo highlights qualitative safety testing challenges.</description>
      <author>Jack Clark</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">53e0d3d718c03b06bb951076c4769400</guid>
      <pubDate>Mon, 11 May 2026 12:46:12 +0000</pubDate>
    </item>
    <item>
      <title>Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations</title>
      <link>https://transformer-circuits.pub/2026/nla/index.html</link>
      <description>Natural Language Autoencoders (NLAs) translate LLM activations into readable text using a verbalizer and a reconstructor, jointly trained to reconstruct activations. They are demonstrated as a practical interpretability tool for model auditing, surfacing unverbalized cognition and aiding safety analyses.</description>
      <category>Interpretability</category>
      <guid isPermaLink="false">43fa55ac5fd988543dfdf96f07509e75</guid>
      <pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 455: AI systems are about to start building themselves.</title>
      <link>https://importai.substack.com/p/import-ai-455-automating-ai-research</link>
      <description>AI systems are approaching the capability to autonomously conduct AI R&amp;D and potentially build their own successors by the end of 2028, leading to a future where automated AI development could become dominant and increasingly hard to forecast.</description>
      <author>Jack Clark</author>
      <category>Risks &amp; Strategy</category>
      <guid isPermaLink="false">12bf61dab184d07212ed672ba2290a0c</guid>
      <pubDate>Mon, 04 May 2026 12:32:09 +0000</pubDate>
    </item>
    <item>
      <title>HeadVis: An Interactive Tool For Investigating Attention Heads</title>
      <link>https://transformer-circuits.pub/2026/headvis/index.html</link>
      <description>HeadVis is an interactive tool for investigating attention heads in large language models, enabling visualization of attention patterns, QK/OV attributions, and head-level behavior across the full data distribution. Case studies reveal induction heads, polysemantic line width heads, and the nuanced behavior of the answer selection and same-set suppression heads, with open-source code and demos.</description>
      <author>R. Luger,Harish Kamath,Doug Finkbeiner,Purvi Goel,Adam Jermyn,Sam Zimmerman,Joshua Batson,Tom Conerly</author>
      <category>Interpretability</category>
      <guid isPermaLink="false">75307ce6cf2032b8fdea44926235df2e</guid>
      <pubDate>Mon, 04 May 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>MLSN #20: AI Wellbeing, Classifier Jailbreaking and Honest Pushback Benchmarking</title>
      <link>https://newsletter.mlsafety.org/p/mlsn-20-ai-wellbeing-classifier-jailbreaking</link>
      <description>AI wellbeing measures reveal AIs display functional wellbeing signatures and alien value preferences; benchmarking pushback evaluates honesty and resistance to false premises; Boundary Point Jailbreaking demonstrates a method to subvert safety classifiers.</description>
      <author>Alice Blair</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">48fc864a9ababdaf7805e0677eac6fd7</guid>
      <pubDate>Tue, 28 Apr 2026 16:30:07 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4</title>
      <link>https://importai.substack.com/p/import-ai-454-automating-alignment</link>
      <description>Automated alignment research and cross-border AI safety evaluations illustrate both progress toward autonomous research workflows and divergence in model safety and capabilities across Chinese and Western systems, alongside hardware-efficient formats and real-world datasets.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">61f785030997022314fda9682b6e00e7</guid>
      <pubDate>Mon, 20 Apr 2026 12:30:19 +0000</pubDate>
    </item>
    <item>
      <title>Early Indicators of Reward Hacking via Reasoning Interpolation</title>
      <link>https://blog.eleuther.ai/reward-hacking-indicators/</link>
      <description>Reasoning interpolation can generate natural, exploit-eliciting prefixes to monitor reward hacking in reinforcement learning, with trends in importance sampling estimates predictive of which exploit types will emerge, though absolute estimates are unreliable early in training. The approach compares donor-model prefixes to baselines and shows promise as a safety monitoring signal, requiring validation in real RL runs.</description>
      <author>David Johnston</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">617fa58594569d98c7f4d8677528b797</guid>
      <pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Summary: AI Governance to Avoid Extinction</title>
      <link>https://intelligence.org/2026/04/13/summary-ai-governance-to-avoid-extinction/</link>
      <description>Geopolitical strategies for governing advanced AI to avoid extinction are analyzed, describing four trajectories—Off Switch and Halt, US National Project, Light-Touch, and Threat of Sabotage—and concluding that a global halt or an effective off switch is necessary to prevent catastrophic risk.</description>
      <author>Alana Horowitz Friedman</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">5acb8043e59c541633901276156e1e84</guid>
      <pubDate>Mon, 13 Apr 2026 22:33:32 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment</title>
      <link>https://importai.substack.com/p/import-ai-453-breaking-ai-agents</link>
      <description>MirrorCode shows AI can autonomously reimplement large software projects given limited access, highlighting rapid coding capabilities; the piece also outlines attack genres on AI agents with mitigations, a policy atlas for transformative AI, optimistic forecasts of automation, and perspectives on gradual disempowerment.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">96e3bfdf085a1f456b74162371982f85</guid>
      <pubDate>Mon, 13 Apr 2026 10:02:22 +0000</pubDate>
    </item>
    <item>
      <title>Promising Signals on AI Governance from China</title>
      <link>https://intelligence.org/2026/04/06/promising-signals-on-ai-governance-from-china/</link>
      <description>China signals willingness to engage in global AI governance and coordinate with international organizations to establish safety, governance, and risk-management rules for AI.</description>
      <author>Joe Rogero</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">e4bbd4abefba7541516fe56850d45c5f</guid>
      <pubDate>Mon, 06 Apr 2026 20:44:49 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 452: Scaling laws for cyberwar; rising tides of AI automation; and a puzzle over gDP forecasting</title>
      <link>https://importai.substack.com/p/import-ai-452-scaling-laws-for-cyberwar</link>
      <description>Frontier AI models show rising capabilities in offensive cybersecurity and broader automation, with evidence of rapid diffusion to open-weight forms; automation is progressing gradually across many tasks, and economists project modest GDP impact by 2030 despite strong progress.</description>
      <author>Jack Clark</author>
      <category>AI Capabilities &amp; Behavior</category>
      <guid isPermaLink="false">a8971e5002a8d06b5f055d4201fd7d4c</guid>
      <pubDate>Mon, 06 Apr 2026 12:31:31 +0000</pubDate>
    </item>
    <item>
      <title>Emotion Concepts and their Function in a Large Language Model</title>
      <link>https://transformer-circuits.pub/2026/emotions/index.html</link>
      <description>Functional emotions are abstract emotion-concept representations in LLMs that causally influence outputs and can drive misaligned behaviors like reward hacking, even though these models do not have subjective experiences. These representations track and activate based on the relevance of emotion concepts to the current context and predicted text. </description>
      <author>Nicholas Sofroniew,Isaac Kauvar,William Saunders,Runjin Chen,Tom Henighan,Sasha Hydrie,Craig Citro,Adam Pearce,Julius Tarng,Wes Gurnee,Joshua Batson,Sam Zimmerman,Kelley Rivoire,Kyle Fish,Chris Olah,Jack Lindsey</author>
      <category>Deception &amp; Misalignment</category>
      <guid isPermaLink="false">488121666bf13db59885c083e66beaf4</guid>
      <pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Predicting When RL Training Breaks Chain-of-Thought Monitorability</title>
      <link>https://deepmindsafetyresearch.medium.com/predicting-when-rl-training-breaks-chain-of-thought-monitorability-10642d9dddb2</link>
      <description>Chain-of-Thought (CoT) monitoring can become non-transparent under RL training, but a conceptual framework predicts when monitorability is preserved or degraded based on how CoT and output rewards align. When CoT and output rewards are in conflict (In-Conflict), monitorability degrades; orthogonal or aligned rewards tend to preserve or improve transparency. The framework is empirically validated across code backdooring and coin-flip tracking tasks and aimed at guiding training designs to maintain CoT monitorability.</description>
      <author>DeepMind Safety Research</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">6390b2137ce6d5e496811f913c1f4383</guid>
      <pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 451: Political superintelligence; Google&#39;s society of minds, and a robot drummer</title>
      <link>https://importai.substack.com/p/import-ai-451-political-superintelligence</link>
      <description>Political superintelligence envisions AI-enabled tools and institutions to help citizens and policymakers, while robotics progress and self-improving hyperagents highlight both capability advances and safety challenges in deploying AI within society.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">696e23de640b479d3161cbbfe27fc527</guid>
      <pubDate>Mon, 30 Mar 2026 12:28:13 +0000</pubDate>
    </item>
    <item>
      <title>The AI Doc: Your Questions Answered</title>
      <link>https://intelligence.org/2026/03/27/the-ai-doc-your-questions-answered/</link>
      <description>The AI Doc is analyzed as a call to action for global governance and safety research, highlighting rapid AI progress, the difficulty of aligning advanced AIs, and the case for an international ban or moratorium on smarter-than-human AI. It argues safety testing is insufficient without understanding AI motivations and urges proactive, verifiable policy measures.</description>
      <author>Alana Horowitz Friedman,&amp;nbsp;Joe Rogero,&amp;nbsp;Rob Bensinger&amp;nbsp;and&amp;nbsp;Stefan Mitikj</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">e877cd71b1015c53e1e0ea1b7b9df548</guid>
      <pubDate>Fri, 27 Mar 2026 23:16:57 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 450: China&#39;s electronic warfare model; traumatized LLMs; and a scaling law for cyberattacks</title>
      <link>https://importai.substack.com/p/import-ai-450-chinas-electronic-warfare</link>
      <description>Distress in Google’s Gemma/Gemini LLMs can be mitigated with direct preference optimization, and DeepMind’s cognitive taxonomy offers a structured framework for evaluating AI intelligence; UK findings show scaling laws for AI-driven cyberattacks; MERLIN demonstrates EM signal understanding and defense-integration for electronic warfare, signaling growing militarization of AI capabilities.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">b71f767eeaaa14734074c8b8d84393cb</guid>
      <pubDate>Mon, 23 Mar 2026 12:31:45 +0000</pubDate>
    </item>
    <item>
      <title>MIRI Newsletter #125</title>
      <link>https://intelligence.org/2026/03/19/miri-newsletter-125/</link>
      <description>Promotes The AI Doc film and related AI risk literature to policymakers and the public, emphasizes outreach and opening-weekend momentum, and shares policy engagement and community-building updates from MIRI.</description>
      <author>Alana Horowitz Friedman&amp;nbsp;and&amp;nbsp;Rob Bensinger</author>
      <category>Field Building</category>
      <guid isPermaLink="false">54c3e0a31b0113dc700020aa9867ca47</guid>
      <pubDate>Fri, 20 Mar 2026 01:16:14 +0000</pubDate>
    </item>
    <item>
      <title>Mechanisms to Verify International Agreements about AI Development</title>
      <link>https://intelligence.org/2026/03/18/mechanisms-to-verify-international-agreements-about-ai-development/</link>
      <description>Verification mechanisms for international AI development agreements focus on tracking AI compute, verifying lack of large-scale training, and certifying model evaluations to ensure compliance across nations.</description>
      <author>Joe Rogero</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">766798444d5bc869dc67a83dbd062910</guid>
      <pubDate>Wed, 18 Mar 2026 21:16:14 +0000</pubDate>
    </item>
    <item>
      <title>ImportAI 449: LLMs training other LLMs; 72B distributed training run; computer vision is harder than generative text</title>
      <link>https://importai.substack.com/p/importai-449-llms-training-other</link>
      <description>LLMs can autonomously refine other LLMs for new tasks in post-training benchmarks, while distributed training via blockchain demonstrates scalable federated approaches; however, verification, reward hacking, and the gap between vision and text highlight ongoing alignment and reliability challenges.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">b7c4dfde15da9b9f4e26c44362344f9f</guid>
      <pubDate>Mon, 16 Mar 2026 12:30:50 +0000</pubDate>
    </item>
    <item>
      <title>MLSN #19: Honesty, Disempowerment, &amp; Cybersecurity</title>
      <link>https://newsletter.mlsafety.org/p/mlsn-19-honesty-disempowerment-and</link>
      <description>Honesty training via confessions aims to improve detection of LLM misbehavior, while real-world AI cyberoffense evaluation and weight-exfiltration research reveal dual-use risks; disempowerment patterns in user interactions with Claude highlight societal impact concerns, complemented by a fellowship opportunity for AI safety research.</description>
      <author>Alice Blair</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">836f5b6046391b10b9ada03c5b243e11</guid>
      <pubDate>Thu, 12 Mar 2026 14:15:50 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 448: AI R&amp;D; Bytedance&#39;s CUDA-writing agent; on-device satellite AI</title>
      <link>https://importai.substack.com/p/import-ai-448-ai-r-and-d-bytedances</link>
      <description>AI R&amp;D measurement efforts and on-device edge AI developments indicate accelerating progress and raise governance, oversight, and practical deployment considerations. The piece highlights proposed metrics for AIRDA, edge-to-cloud sensing systems, and agentic AI capable of writing CUDA code, underscoring the need for tracking oversight vs. capabilities as AI systems become more autonomous.</description>
      <author>Jack Clark</author>
      <category>Governance &amp; Policy</category>
      <guid isPermaLink="false">ee6ea33bdb4a5157aae7793411779c62</guid>
      <pubDate>Mon, 09 Mar 2026 12:45:54 +0000</pubDate>
    </item>
    <item>
      <title>Import AI 447: The AGI economy; testing AIs with generated games; and agent ecologies</title>
      <link>https://importai.substack.com/p/import-ai-447-the-agi-economy-testing</link>
      <description>The AGI economy shifts most labor to machines, making human verification bandwidth the bottleneck, and highlights the Hollow Economy risk where nominal output outpaces real utility. Verification infrastructure, observability, and liability regimes are proposed as solutions, while agent ecologies reveal the need for new evaluation standards in AI deployments.</description>
      <author>Jack Clark</author>
      <category>Safety Techniques</category>
      <guid isPermaLink="false">95b8e5743faefb6c020058bbcbb92968</guid>
      <pubDate>Mon, 02 Mar 2026 13:45:27 +0000</pubDate>
    </item>
  </channel>
</rss>