arXiv:2506.19732cs.LGcs.AI2025-06被引 1

用博弈论方法量化神经元对输出的贡献,揭示模型内部功能分工。

Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units

  • 基于多重扰动的谢尔普利值分析,精准计算每个神经元的贡献。
  • 在560亿参数模型中发现计算集中在少数核心神经元,语言专家具特定分工。
  • 适用于模型解释、编辑与压缩,适合研究者与工程师深入理解大模型。

深度神经网络如今以数十亿参数生成文本、图像和语音,亟需明确每个神经元对高维输出的贡献。现有可解释AI方法如SHAP仅能归因输入,无法量化数千个输出像素、词元或逻辑值中的神经元贡献。本文提出多扰动谢尔普利值分析(MSA),一种模型无关的博弈论框架。通过系统性地破坏神经元组合,MSA生成与模型输出维度一致的谢尔普利模式,得到单元级贡献图。该方法应用于从多层感知机到560亿参数的Mixtral-8x7B及生成对抗网络(GAN)。结果表明,正则化使计算集中于少数枢纽神经元;在大型语言模型中识别出语言特异性专家;在GAN中揭示倒置的像素生成层级结构。这些成果展示了MSA在解释、编辑与压缩深度神经网络方面的强大能力。

原文摘要 · Abstract (English)

Neural networks now generate text, images, and speech with billions of parameters, producing a need to know how each neural unit contributes to these high-dimensional outputs. Existing explainable-AI methods, such as SHAP, attribute importance to inputs, but cannot quantify the contributions of neural units across thousands of output pixels, tokens, or logits. Here we close that gap with Multiperturbation Shapley-value Analysis (MSA), a model-agnostic game-theoretic framework. By systematically lesioning combinations of units, MSA yields Shapley Modes, unit-wise contribution maps that share the exact dimensionality of the model's output. We apply MSA across scales, from multi-layer perceptrons to the 56-billion-parameter Mixtral-8x7B and Generative Adversarial Networks (GAN). The approach demonstrates how regularisation concentrates computation in a few hubs, exposes language-specific experts inside the LLM, and reveals an inverted pixel-generation hierarchy in GANs. Together, these results showcase MSA as a powerful approach for interpreting, editing, and compressing deep neural networks.

模型解释神经元贡献博弈论大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。