arXiv:2506.18852cs.CLcs.AI2025-06被引 9

MI研究需哲学加持,厘清概念、优化方法、应对伦理挑战

Mechanistic Interpretability Needs Philosophy

  • 引入哲学反思机制,梳理MI中的隐含假设与解释策略
  • 通过三个开放问题展示哲学对方法论的改进价值
  • 适合关注AI解释性与伦理的跨学科研究者阅读

机制可解释性(Mechanistic Interpretability, MI)旨在通过揭示神经网络的内在机制来解释其工作原理。随着该领域影响力扩大,不仅需要审视模型本身,还需反思MI研究中隐含的假设、概念和解释策略。本文主张,机制可解释性亟需哲学作为持续合作伙伴,以澄清概念、优化方法,并应对解释人工智能系统时的认知与伦理复杂性。通过分析MI文献中的三个开放问题,本文展示了哲学在推动进展方面的未被充分发掘潜力,并提出了深化跨学科对话的路径。

原文摘要 · Abstract (English)

Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.

机制可解释性哲学AI伦理跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。