arXiv:2510.09135cs.CVcs.LG2025-10被引 3

将视觉模型的预测归因到训练图像的具体区域,揭示模型内部机制。

Training Feature Attribution for Vision Models

  • 从训练图像中定位影响预测的关键区域,实现细粒度解释。
  • 能识别导致误分类的有害训练样本和虚假关联模式。
  • 适合研究模型可解释性与鲁棒性的研究人员使用。

深度神经网络常被视为黑箱,亟需可解释性方法以增强信任与问责。现有方法通常将测试时的预测归因于输入特征(如图像像素)或有影响力的训练样本。本文主张应联合考察两者。本工作提出训练特征归因,将测试预测与特定训练图像的特定区域关联,从而揭示深度模型的内在工作机制。在视觉数据集上的实验表明,该方法生成细粒度、针对具体测试样本的解释:可识别导致误分类的有害样本,并揭示传统归因方法无法发现的虚假关联,如基于补丁的捷径学习模式。

原文摘要 · Abstract (English)

Deep neural networks are often considered opaque systems, prompting the need for explainability methods to improve trust and accountability. Existing approaches typically attribute test-time predictions either to input features (e.g., pixels in an image) or to influential training examples. We argue that both perspectives should be studied jointly. This work explores *training feature attribution*, which links test predictions to specific regions of specific training images and thereby provides new insights into the inner workings of deep models. Our experiments on vision datasets show that training feature attribution yields fine-grained, test-specific explanations: it identifies harmful examples that drive misclassifications and reveals spurious correlations, such as patch-based shortcuts, that conventional attribution methods fail to expose.

可解释性视觉模型特征归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。