arXiv:2412.19523cs.AIcs.CV2024-12

通过对抗性迁移提升模型解释精度,显著增强可解释性。

Attribution for Enhanced Explanation with Transferable Adversarial eXploration

  • 引入可迁移对抗攻击方法,优化归因生成。
  • 在ImageNet上平均比基线高7.57%,优于32.62%现有方法。
  • 适用于各类视觉模型,解释结果更稳定可靠。

深度神经网络的可解释性对理解其决策至关重要,尤其在计算机视觉领域。基于AttEXplore的AttEXplore++框架,引入MIG和GRA等可迁移对抗攻击方法,显著提升解释的准确性和鲁棒性。我们在五种模型(Inception-v3、ResNet-50、VGG16、MaxViT-T、ViT-B/16)上使用ImageNet数据集进行大量实验,结果表明该方法平均性能较AttEXplore提升7.57%,较其他先进解释算法高出32.62%。通过插入与删除评分评估,证实对抗可迁移性在增强归因效果中起关键作用。我们还研究了随机性、扰动率、噪声幅度及多样性概率对归因的影响,证明AttEXplore++在不同模型上均能提供更稳定可靠的解释。代码已公开:https://anonymous.4open.science/r/ATTEXPLOREP-8435/

原文摘要 · Abstract (English)

The interpretability of deep neural networks is crucial for understanding model decisions in various applications, including computer vision. AttEXplore++, an advanced framework built upon AttEXplore, enhances attribution by incorporating transferable adversarial attack methods such as MIG and GRA, significantly improving the accuracy and robustness of model explanations. We conduct extensive experiments on five models, including CNNs (Inception-v3, ResNet-50, VGG16) and vision transformers (MaxViT-T, ViT-B/16), using the ImageNet dataset. Our method achieves an average performance improvement of 7.57\% over AttEXplore and 32.62\% compared to other state-of-the-art interpretability algorithms. Using insertion and deletion scores as evaluation metrics, we show that adversarial transferability plays a vital role in enhancing attribution results. Furthermore, we explore the impact of randomness, perturbation rate, noise amplitude, and diversity probability on attribution performance, demonstrating that AttEXplore++ provides more stable and reliable explanations across various models. We release our code at: https://anonymous.4open.science/r/ATTEXPLOREP-8435/

可解释性对抗攻击归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。