arXiv:2601.00968cs.LG2026-01被引 5

用解释性指导模型优化,提升对抗攻击防御能力。

Explainability-Guided Defense: Attribution-Aware Model Refinement Against Adversarial Data Attacks

  • 利用LIME识别干扰特征,通过掩码与正则化抑制其影响
  • 在CIFAR-10/100上实现更强的抗干扰与泛化性能
  • 无需额外数据或架构,可融入标准对抗训练流程

深度学习在医疗、自动驾驶等关键领域应用日益广泛,亟需兼具鲁棒性与可解释性的防御机制。本文发现,通过局部可解释模型无关解释(LIME)识别出的虚假、不稳定或语义无关特征,会显著加剧模型对对抗扰动的脆弱性。基于此,我们提出一种归因引导的模型精炼框架,将LIME从被动诊断工具转变为主动训练信号。方法通过特征掩码、感知敏感度正则化与对抗增强,在闭环中系统抑制干扰特征。该策略不依赖额外数据或模型结构,可无缝集成至标准对抗训练。理论上,我们推导出归因一致性与鲁棒性之间的下界关系。在CIFAR-10、CIFAR-10-C和CIFAR-100上的实证表明,该方法显著提升了对抗鲁棒性与分布外泛化能力。

原文摘要 · Abstract (English)

The growing reliance on deep learning models in safety-critical domains such as healthcare and autonomous navigation underscores the need for defenses that are both robust to adversarial perturbations and transparent in their decision-making. In this paper, we identify a connection between interpretability and robustness that can be directly leveraged during training. Specifically, we observe that spurious, unstable, or semantically irrelevant features identified through Local Interpretable Model-Agnostic Explanations (LIME) contribute disproportionately to adversarial vulnerability. Building on this insight, we introduce an attribution-guided refinement framework that transforms LIME from a passive diagnostic into an active training signal. Our method systematically suppresses spurious features using feature masking, sensitivity-aware regularization, and adversarial augmentation in a closed-loop refinement pipeline. This approach does not require additional datasets or model architectures and integrates seamlessly into standard adversarial training. Theoretically, we derive an attribution-aware lower bound on adversarial distortion that formalizes the link between explanation alignment and robustness. Empirical evaluations on CIFAR-10, CIFAR-10-C, and CIFAR-100 demonstrate substantial improvements in adversarial robustness and out-of-distribution generalization.

对抗防御可解释性模型精炼LIME

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。