用机制解释法破解深度学习黑箱,让AI决策更透明
A Mechanistic Explanatory Strategy for XAI
- 从科学哲学出发,用机制分解法解析神经网络决策过程
- 案例验证可发现传统方法忽略的功能组件与激活模式
- 适合研究AI可解释性与机制推理的学者和工程师
尽管可解释人工智能(XAI)取得进展,学界仍普遍认为其缺乏稳固的概念基础,并未融入更广泛的科学解释话语体系。为此,本文提出一种基于科学哲学中机制解释策略的框架,用于揭示深度学习系统的功能组织。该方法强调识别驱动决策的功能机制,包括神经元、层、电路或激活模式等关键组件,并通过分解、定位与重构来理解其作用。图像识别与语言建模的初步案例研究表明,这一策略与OpenAI、Anthropic的机制可解释性研究相契合。结果表明,采用机制解释能发现传统技术遗漏的重要元素,推动实现更全面的可解释人工智能。
原文摘要 · Abstract (English)
Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific discourse on explanation. In response, emerging research draws on explanatory strategies from various sciences and the philosophy of science literature to fill these gaps. This paper outlines a mechanistic strategy for explaining the functional organization of deep learning systems, situating recent developments in explainable AI within a broader philosophical context. According to the mechanistic approach, the explanation of opaque AI systems involves identifying mechanisms that drive decision making. For deep neural networks, this means discerning functionally relevant components, such as neurons, layers, circuits, or activation patterns, and understanding their roles through decomposition, localization, and recomposition. Proof-of-principle case studies from image recognition and language modeling align these theoretical approaches with mechanistic interpretability research from OpenAI and Anthropic. The findings suggest that pursuing mechanistic explanations can uncover elements that traditional explainability techniques may overlook, ultimately contributing to more thoroughly explainable AI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。