arXiv:2411.03993cs.CV2024-11被引 5

对比神经元与字典基表示,发现后者更易被人类理解。

Choosing the right basis for interpretability: Psychophysical comparison between neuron-based and dictionary-based representations

  • 用字典学习法重构网络层激活,获得更清晰的解释基
  • 深度网络中字典基解释力比神经元基强,深层优势更明显
  • 建议用字典基评估模型可解释性,避免神经元基误导

可解释性研究常以单个神经元为基本单位,但神经元存在叠加效应,导致解释不清晰。字典学习方法(如稀疏自编码器、非负矩阵分解)通过学习新基来替代神经元,提供更优解释。我们开展了三项大规模在线心理物理学实验(共481人),在两个卷积网络(ResNet50、VGG16)中比较基于神经元和字典基的解释效果。通过视觉连贯性衡量可解释性:若人类能稳定识别最大激活图像中的共同视觉模式并推广至新图像,则该基更可解释。结果表明,字典基始终优于神经元基,且在深层网络中优势更显著。值得注意的是,由于模型间神经元对齐程度不同(如ResNet50叠加更严重),仅用神经元分析会掩盖跨模型差异——其更高可解释性仅在字典基下显现。这些结果为字典基作为可解释性基础提供了心理学证据,并警示仅依赖神经元分析可能产生偏差。

原文摘要 · Abstract (English)

Interpretability research often adopts a neuron-centric lens, treating individual neurons as the fundamental units of explanation. However, neuron-level explanations can be undermined by superposition, where single units respond to mixtures of unrelated patterns. Dictionary learning methods, such as sparse autoencoders and non-negative matrix factorization, offer a promising alternative by learning a new basis over layer activations. Despite this promise, direct human evaluations comparing neuron-based and dictionary-based representations remain limited. We conducted three large-scale online psychophysics experiments (N=481) comparing explanations derived from neuron-based and dictionary-based representations in two convolutional neural networks (ResNet50, VGG16). We operationalize interpretability via visual coherence: a basis is more interpretable if humans can reliably recognize a common visual pattern in its maximally activating images and generalize that pattern to new images. Across experiments, dictionary-based representations were consistently more interpretable than neuron-based representations, with the advantage increasing in deeper layers. Critically, because models differ in how neuron-aligned their representations are -- with ResNet50 exhibiting greater superposition, neuron-based evaluations can mask cross-model differences, such that ResNet50's higher interpretability emerges only under dictionary-based comparisons. These results provide psychophysical evidence that dictionary-based representations offer a stronger foundation for interpretability and caution against model comparisons based solely on neuron-level analyses.

可解释性字典学习神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。