arXiv:2511.19265cs.LG2025-11被引 6

解析神经网络内部计算机制,让黑箱模型可理解

Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks

  • 通过逆向工程揭示神经网络的内在算法逻辑
  • 提出统一分类框架,系统梳理关键解释技术
  • 适合想深入理解模型原理的研究者入门

深度神经网络的黑箱特性给透明可信的人工智能部署带来挑战。随着人工智能在社会中日益普及,开发能够解释和理解其决策的方法变得至关重要。为此,机制可解释性(MI)作为可解释人工智能(XAI)领域的一个有前景且独特的研究方向应运而生。MI是研究神经网络内部计算过程,并将其转化为人类可理解算法的过程,涵盖旨在揭示神经网络实现的计算算法的逆向工程方法。本文提出一个统一的MI方法分类体系,详细分析关键技术,辅以具体案例和伪代码说明。将MI置于更广阔的可解释性图景中,对比其目标、方法与洞察与其他XAI流派的异同。同时追溯了MI作为研究领域的演进历程,强调其概念根源及近年研究的加速发展。我们认为,MI具有推动机器学习系统科学理解的巨大潜力——将模型不仅视为任务求解工具,更视为可研究与理解的系统。我们希望吸引新研究者进入机制可解释性领域。

原文摘要 · Abstract (English)

The black box nature of deep neural networks poses a significant challenge for the deployment of transparent and trustworthy artificial intelligence (AI) systems. With the growing presence of AI in society, it becomes increasingly important to develop methods that can explain and interpret the decisions made by these systems. To address this, mechanistic interpretability (MI) emerged as a promising and distinctive research program within the broader field of explainable artificial intelligence (XAI). MI is the process of studying the inner computations of neural networks and translating them into human-understandable algorithms. It encompasses reverse engineering techniques aimed at uncovering the computational algorithms implemented by neural networks. In this article, we propose a unified taxonomy of MI approaches and provide a detailed analysis of key techniques, illustrated with concrete examples and pseudo-code. We contextualize MI within the broader interpretability landscape, comparing its goals, methods, and insights to other strands of XAI. Additionally, we trace the development of MI as a research area, highlighting its conceptual roots and the accelerating pace of recent work. We argue that MI holds significant potential to support a more scientific understanding of machine learning systems -- treating models not only as tools for solving tasks, but also as systems to be studied and understood. We hope to invite new researchers into the field of mechanistic interpretability.

机制可解释性神经网络XAI逆向工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。