arXiv:2503.00641cs.CV2025-03ICLR被引 9

模型分类层训练细节比预训练方法更影响解释质量

How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations

  • 调整分类层可显著提升后验解释效果
  • 分类层参数仅占模型10%却主导解释结果
  • 适合关注模型可解释性的研究人员

后验重要性归因方法常被用于解释深度神经网络,其假设解释可独立于模型训练过程。然而本文通过实证发现,这一假设不成立:预训练模型的分类层(参数占比不足10%)训练细节对解释质量的影响远超预训练方案本身。该发现具有重要实践意义,尤其在预训练技术日益多样化的背景下。我们提出简单有效的分类层调整方法,可在多种视觉预训练框架(全监督、自监督、对比视觉语言训练)中显著提升解释质量,并在多种评估指标下验证了其有效性。

原文摘要 · Abstract (English)

Post-hoc importance attribution methods are a popular tool for "explaining" Deep Neural Networks (DNNs) and are inherently based on the assumption that the explanations can be applied independently of how the models were trained. Contrarily, in this work we bring forward empirical evidence that challenges this very notion. Surprisingly, we discover a strong dependency on and demonstrate that the training details of a pre-trained model's classification layer (less than 10 percent of model parameters) play a crucial role, much more than the pre-training scheme itself. This is of high practical relevance: (1) as techniques for pre-training models are becoming increasingly diverse, understanding the interplay between these techniques and attribution methods is critical; (2) it sheds light on an important yet overlooked assumption of post-hoc attribution methods which can drastically impact model explanations and how they are interpreted eventually. With this finding we also present simple yet effective adjustments to the classification layers, that can significantly enhance the quality of model explanations. We validate our findings across several visual pre-training frameworks (fully-supervised, self-supervised, contrastive vision-language training) and analyse how they impact explanations for a wide range of attribution methods on a diverse set of evaluation metrics.

可解释性后验解释分类层DNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。