发现并还原深度视觉模型中概念的编码与解码方向对,揭示模型内部工作机制。
Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks
- 通过激活聚类和概率信号向量,无监督地识别概念的解码与编码方向。
- 在真实数据上解码方向对应可解释语义概念,性能优于现有无监督方法。
- 可应用于模型解释、预测分析及错误修正,提升模型可信赖性。
实证表明,深度视觉网络常将概念表示为潜在空间中的方向,概念信息以向量分量形式写入输入表示。然而,编码(写入)与解码(读取)机制是训练过程自然涌现的隐含机制,难以直接获取。恢复该机制有助于打开深度学习模型的黑箱,实现理解、调试与改进。本文提出一种无监督方法:在假设线性概念表示的前提下,每个概念可通过一对方向实现编码与解码——前者促进信息写入,后者支持信息读取。不同于依赖特征重建的矩阵分解、自编码器或字典学习方法,本工作采用新视角:通过激活的方向聚类识别解码方向,基于概率视角估计编码信号向量。我们进一步引入新型技术‘不确定性区域对齐’,利用网络权重揭示影响预测的可解释方向。分析显示:(a) 在合成数据上,方法能准确恢复真实方向对;(b) 在真实数据上,解码方向映射至单义、可解释概念,且优于无监督基线;(c) 信号向量能忠实估计编码方向,经激活最大化验证。最后,展示了其在理解全局模型行为、解释个体预测、生成反事实或纠正错误方面的应用。
原文摘要 · Abstract (English)
Empirical evidence shows that deep vision networks often represent concepts as directions in latent space with concept information written along directional components in the vector representation of the input. However, the mechanism to encode (write) and decode (read) concept information to and from vector representations is not directly accessible as it constitutes a latent mechanism that naturally emerges from the training process of the network. Recovering this mechanism unlocks significant potential to open the black-box nature of deep networks, enabling understanding, debugging, and improving deep learning models. In this work, we propose an unsupervised method to recover this mechanism. For each concept, we explain that under the hypothesis of linear concept representations, this mechanism can be implemented with the help of two directions: the first facilitating encoding of concept information and the second facilitating decoding. Unlike prior matrix decomposition, autoencoder, or dictionary learning methods that rely on feature reconstruction, we propose a new perspective: decoding directions are identified via directional clustering of activations, and encoding directions are estimated with signal vectors under a probabilistic view. We further leverage network weights through a novel technique, Uncertainty Region Alignment, which reveals interpretable directions affecting predictions. Our analysis shows that (a) on synthetic data, our method recovers ground-truth direction pairs; (b) on real data, decoding directions map to monosemantic, interpretable concepts and outperform unsupervised baselines; and (c) signal vectors faithfully estimate encoding directions, validated via activation maximization. Finally, we demonstrate applications in understanding global model behavior, explaining individual predictions, and intervening to produce counterfactuals or correct errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。