arXiv:2504.11489cs.CV2025-04CVPR被引 1

用稀疏自编码器揭示InceptionV1中分支的专化现象

Uncovering Branch specialization in InceptionV1 using k sparse autoencoders

  • 通过稀疏自编码器分析InceptionV1各层分支的特征专化
  • 发现mixed4a-4e等分支中存在稳定一致的特征分布模式
  • 适合研究神经网络可解释性与结构设计的学者参考

稀疏自编码器(SAEs)在从神经网络中提取可解释特征方面表现优异,尤其能缓解由叠加导致的多义神经元问题。已有研究表明,SAEs可有效解析InceptionV1浅层的特征。尽管此后对SAE进行了多项改进,但InceptionV1深层中分支的专化现象仍不明确。本文展示了mixed4a-4e分支、5x5分支及一个1x1分支中各层均存在分支专化现象。同时,我们提供证据表明,相似特征在模型不同层中会稳定地定位在相同卷积尺寸的分支中,说明分支专化具有跨层一致性。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have shown to find interpretable features in neural networks from polysemantic neurons caused by superposition. Previous work has shown SAEs are an effective tool to extract interpretable features from the early layers of InceptionV1. Since then, there have been many improvements to SAEs but branch specialization is still an enigma in the later layers of InceptionV1. We show various examples of branch specialization occuring in each layer of the mixed4a-4e branch, in the 5x5 branch and in one 1x1 branch. We also provide evidence to claim that branch specialization seems to be consistent across layers, similar features across the model will be localized in the same convolution size branches in their respective layer.

神经网络解释分支专化稀疏自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。