arXiv:2509.21387cs.CVcs.AI2025-09被引 1

轻度剪枝可让模型注意力更贴近人类认知,过度剪枝则损害可解释性。

Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence

  • 通过权重裁剪+微调,观察模型注意力图变化
  • 轻度剪枝提升注意力聚焦度与语义一致性,但过量剪枝会破坏特征独立性
  • 适合关注模型可解释性与剪枝平衡的研究者

先前研究显示神经网络可在大幅裁剪后仍保持性能,但剪枝对模型可解释性的影响尚不明确。本文以在ImageNette上训练的ResNet-18为对象,研究基于幅度的剪枝结合微调对低层显著性图与高层概念表征的影响。比较了不同剪枝程度下原始梯度(VG)与积分梯度(IG)的后验解释,评估其稀疏性与忠实性。进一步采用CRAFT方法追踪学习到的概念语义一致性变化。结果显示,轻度至中度剪枝能提升显著性图的聚焦度与忠实性,同时保持清晰、语义明确的概念;而激进剪枝虽维持准确率,却导致异质特征融合,降低显著性图稀疏性与概念一致性。表明剪枝可引导内部表示趋向更符合人类注意力模式,但过度剪枝会削弱可解释性。

原文摘要 · Abstract (English)

Prior works have shown that neural networks can be heavily pruned while preserving performance, but the impact of pruning on model interpretability remains unclear. In this work, we investigate how magnitude-based pruning followed by fine-tuning affects both low-level saliency maps and high-level concept representations. Using a ResNet-18 trained on ImageNette, we compare post-hoc explanations from Vanilla Gradients (VG) and Integrated Gradients (IG) across pruning levels, evaluating sparsity and faithfulness. We further apply CRAFT-based concept extraction to track changes in semantic coherence of learned concepts. Our results show that light-to-moderate pruning improves saliency-map focus and faithfulness while retaining distinct, semantically meaningful concepts. In contrast, aggressive pruning merges heterogeneous features, reducing saliency map sparsity and concept coherence despite maintaining accuracy. These findings suggest that while pruning can shape internal representations toward more human-aligned attention patterns, excessive pruning undermines interpretability.

模型剪枝可解释性注意力机制概念一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。