验证扩散去噪平滑提升视觉Transformer解释性鲁棒性
[Re] Improving Interpretation Faithfulness for Vision Transformers
- 用扩散去噪平滑增强解释方法对攻击和扰动的抗性
- 实测表明该方法在分类与分割任务中均提升解释鲁棒性
- 适用于关注模型可解释性安全的研究者
本研究旨在复现arXiv:2311.17983提出的忠实视觉Transformer(FViT)及其解释方法,包括arXiv:2012.09838和Xu(2022)等提出的方案。我们检验了arXiv:2311.17983中的核心主张:使用扩散去噪平滑(DDS)能提升解释性在(1)分割任务中的抗攻击鲁棒性,以及(2)分类任务中对扰动和攻击的鲁棒性。同时扩展原研究,测试将DDS应用于任意解释方法是否均可提升其抗攻击能力,涵盖基线方法及近期提出的归因传播法(Attribution Rollout)。此外,还评估了通过DDS构建FViT的计算成本与环境影响。结果总体支持原研究结论,但发现若干细微差异并予以讨论。
原文摘要 · Abstract (English)
This work aims to reproduce the results of Faithful Vision Transformers (FViTs) proposed by arXiv:2311.17983 alongside interpretability methods for Vision Transformers from arXiv:2012.09838 and Xu (2022) et al. We investigate claims made by arXiv:2311.17983, namely that the usage of Diffusion Denoised Smoothing (DDS) improves interpretability robustness to (1) attacks in a segmentation task and (2) perturbation and attacks in a classification task. We also extend the original study by investigating the authors' claims that adding DDS to any interpretability method can improve its robustness under attack. This is tested on baseline methods and the recently proposed Attribution Rollout method. In addition, we measure the computational costs and environmental impact of obtaining an FViT through DDS. Our results broadly agree with the original study's findings, although minor discrepancies were found and discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。