用人类理解自身的方式解释AI,让机器的决策像人一样可被理解。
Propositional Interpretability in Artificial Intelligence
- 以信念、欲望等命题态度解释AI内部机制
- 提出思想日志概念,记录系统随时间的命题态度变化
- 适合关注AI可解释性与哲学基础的研究者
机械可解释性旨在通过系统内部机制解释AI的行为。本文分析该领域的若干方面,指出当前挑战并评估进展。强调命题可解释性的重要性,即通过命题态度(如信念、欲望或主观概率)来解释系统机制与行为。命题态度是人类自我理解的核心方式,也可能是未来AI解释的关键。核心挑战在于‘思想日志’:构建能记录AI系统在时间维度上所有相关命题态度的系统。文章评估了当前主流可解释性方法(如探测、稀疏自编码器、思维链)及基于心理语义学的哲学解释方法在实现命题可解释性上的优劣。
原文摘要 · Abstract (English)
Mechanistic interpretability is the program of explaining what AI systems are doing in terms of their internal mechanisms. I analyze some aspects of the program, along with setting out some concrete challenges and assessing progress to date. I argue for the importance of propositional interpretability, which involves interpreting a system's mechanisms and behavior in terms of propositional attitudes: attitudes (such as belief, desire, or subjective probability) to propositions (e.g. the proposition that it is hot outside). Propositional attitudes are the central way that we interpret and explain human beings and they are likely to be central in AI too. A central challenge is what I call thought logging: creating systems that log all of the relevant propositional attitudes in an AI system over time. I examine currently popular methods of interpretability (such as probing, sparse auto-encoders, and chain of thought methods) as well as philosophical methods of interpretation (including those grounded in psychosemantics) to assess their strengths and weaknesses as methods of propositional interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。