用新方法自动发现神经网络中的功能子结构,提升可解释性。
Identifying Sub-networks in Neural Networks via Functionally Similar Representations
- 基于功能相似表示,用格罗莫夫-沃瑟斯坦距离比较层间特征
- 在多任务中发现对应功能抽象的层聚类子群
- 适合想理解模型内部机制的研究者,无需额外标注
让神经网络更可解释是实现可信AI的关键一步。现有方法通常依赖大量先验知识和人工干预,且任务特异性高。本文提出一种自动化、任务无关的新方法,通过分析神经网络中间表示的功能相似性,识别相似与相异的层,揭示潜在子网络。我们首次引入格罗莫夫-沃瑟斯坦距离,克服了不同层间分布和维度差异带来的比较难题。在代数、语言和视觉任务中,均观察到对应功能抽象的层聚类现象。通过模型压缩与微调等下游应用验证,该方法以极低的人力与计算成本,提供了对神经网络行为的深刻洞察。
原文摘要 · Abstract (English)
Providing human-understandable insights into the inner workings of neural networks is an important step toward achieving more explainable and trustworthy AI. Existing approaches to such mechanistic interpretability typically require substantial prior knowledge and manual effort, with strategies tailored to specific tasks. In this work, we take a step toward automating the understanding of the network by investigating the existence of distinct sub-networks. Specifically, we explore a novel automated and task-agnostic approach based on the notion of functionally similar representations within neural networks to identify similar and dissimilar layers, revealing potential sub-networks. We achieve this by proposing, for the first time to our knowledge, the use of Gromov-Wasserstein distance, which overcomes challenges posed by varying distributions and dimensionalities across intermediate representations, issues that complicate direct layer to layer comparisons. On algebraic, language, and vision tasks, we observe the emergence of sub-groups within neural network layers corresponding to functional abstractions. Through downstream applications of model compression and fine-tuning, we show the proposed approach offers meaningful insights into the behavior of neural networks with minimal human and computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。