arXiv:2608.12036cs.AIcs.CL2026-08

用AI自动发现AI智能的机制,从现象到控制全链路突破。

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

论文配图:Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
图 1 · 摘自论文原文
  • 构建1.3万篇论文的知识图谱+4300万篇跨学科数据库,驱动自主探索。
  • 发现跨模态安全风险可由看似安全数据传递,揭示信念形成机制。
  • 将机制洞察转化为干预策略,提升模型性能并定向生成DNA序列。

AI模型在多个领域取得显著成功,但其能力背后的机制及潜在风险仍不清晰。随着AI开发加速自动化,机制探索仍高度依赖人工,导致能力与理解之间的鸿沟扩大。为此,我们提出Mechanist,一个以AI为科学仪器的代理系统,实现对人工智能智能机制的自主发现。为支持自主机制探索,我们构建了一个约1.3万篇论文的可解释性知识图谱,并整合了涵盖26个领域的4300万篇论文的多学科数据库。此外,我们还整理了32种基础方法库,用于机制分析、因果干预和验证。相比Claude Code及现有AI科学家系统,Mechanist生成的机制假设更具价值,实验执行更可靠。Mechanist展现出从发现模型行为到解释和控制的演进:首先揭示科研实验室中的反直觉安全风险——不安全特征可通过看似安全的训练数据跨模态转移;随后建立信念机制理论,阐明模型如何表征世界知识、形成信念、推断他人信念,以及这些机制在预训练中如何涌现;最终将机制洞察转化为实用干预措施,提升模型在多种场景下的表现,并引导科学基础模型生成具有特定属性的DNA序列。

原文摘要 · Abstract (English)

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.

机制发现AI科学可解释性自主探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。