arXiv:2412.18952cs.LGcs.AI2024-12被引 14

用LIME指导模型优化,让模型更可靠且更容易理解。

Bridging Interpretability and Robustness Using LIME-Guided Model Refinement

  • 用LIME识别无关特征,训练时惩罚模型对这些特征的依赖。
  • 在多个数据集上提升对抗样本抵抗能力和分布外泛化性能。
  • 适合关注模型可解释性与鲁棒性平衡的研究者使用。

本文探讨深度学习模型中可解释性与鲁棒性之间的复杂关系。尽管在各类任务中表现优异,深度模型仍存在易受对抗攻击、过度依赖虚假相关性及决策过程不透明等关键缺陷。为此,本文提出一种新框架,利用局部可解释模型无关解释(LIME)系统性提升模型鲁棒性。通过识别并缓解无关或误导性特征的影响,该方法在训练中迭代优化模型,抑制对这些特征的依赖。在多个基准数据集上的实证评估表明,基于LIME的模型精炼不仅提升了可解释性,还显著增强了对抗扰动的抵抗能力,并改善了对分布外数据的泛化性能。

原文摘要 · Abstract (English)

This paper explores the intricate relationship between interpretability and robustness in deep learning models. Despite their remarkable performance across various tasks, deep learning models often exhibit critical vulnerabilities, including susceptibility to adversarial attacks, over-reliance on spurious correlations, and a lack of transparency in their decision-making processes. To address these limitations, we propose a novel framework that leverages Local Interpretable Model-Agnostic Explanations (LIME) to systematically enhance model robustness. By identifying and mitigating the influence of irrelevant or misleading features, our approach iteratively refines the model, penalizing reliance on these features during training. Empirical evaluations on multiple benchmark datasets demonstrate that LIME-guided refinement not only improves interpretability but also significantly enhances resistance to adversarial perturbations and generalization to out-of-distribution data.

可解释性鲁棒性LIME模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。