arXiv:2503.09735cs.CRcs.CV2025-03

用模型解释提升对抗样本检测,但现有方法受环境依赖性强。

Enhancing Adversarial Example Detection Through Model Explanation

  • 通过模型解释分析输入异常,识别对抗样本
  • 原方法在不同环境下性能波动大,鲁棒性不足
  • 适合关注防御系统稳定性的研究者参考

对抗样本是机器学习模型面临的主要问题,促使持续探索有效防御手段。一种有前景的方向是利用模型解释来深入理解并防御此类攻击。我们研究了由NeurIPS 2018亮点论文提出的AmI方法,该方法借助模型解释检测对抗样本。研究表明,尽管AmI思路具有潜力,但其性能高度依赖特定设置(如超参数)以及操作系统和深度学习框架等外部因素,这些缺陷限制了其实际应用。我们的发现强调了开发在多种条件下仍有效的鲁棒防御机制的必要性,并呼吁建立更全面的防御技术评估框架。

原文摘要 · Abstract (English)

Adversarial examples are a major problem for machine learning models, leading to a continuous search for effective defenses. One promising direction is to leverage model explanations to better understand and defend against these attacks. We looked at AmI, a method proposed by a NeurIPS 2018 spotlight paper that uses model explanations to detect adversarial examples. Our study shows that while AmI is a promising idea, its performance is too dependent on specific settings (e.g., hyperparameter) and external factors such as the operating system and the deep learning framework used, and such drawbacks limit AmI's practical usage. Our findings highlight the need for more robust defense mechanisms that are effective under various conditions. In addition, we advocate for a comprehensive evaluation framework for defense techniques.

对抗样本模型解释防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。