arXiv:2606.26523cs.AIcs.LG2026-06

用哲学方法解析AI的信念与意图,提升系统可信赖度。

Radical AI Interpretability

  • 将哲学中的激进解释理论用于分析AI内部认知结构
  • 提出联合约束的信念欲望判定标准,避免片面解读
  • 适合关注AI安全与可信推理的研究者阅读

我们构建了一个将AI系统视为代理的解释框架,结合激进解释哲学传统与机制可解释性工具。核心问题是:在已知系统计算事实的前提下,如何推断其信念、欲望与意义?这对安全性日益重要。我们希望信任所部署的系统,要么理解其目标,要么至少可靠检测欺骗行为。尽管可解释性研究正在开发从模型内部读取信念与欲望的工具,但尚未有明确标准判断这些工具是否成功。本文提出代表主义与解释主义两种路径的评判标准,并将其与现有可解释性方法可执行的测试相联系。关键洞见是:这些属性不能逐个单独推断。信念、欲望及其预设的命题结构是相互约束的整体;一种只固定某一属性而测量其他属性的方法,会继承由此引入的扭曲。这种整体性对不共享解释者概念的AI系统尤为紧迫,但也提供了突破口:系统的态度限制其命题结构,该结构又反过来限制可归因的态度,而机制可解释性可帮助我们量化二者。

原文摘要 · Abstract (English)

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.

AI可解释性信念推理系统安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。