保护关键图像区域,让模型既准确又抗干扰。
Salient Information Preserving Adversarial Training Improves Clean and Robust Accuracy
- 用显著区域引导对抗训练,保留重要特征
- 在多个数据集上提升干净准确率,同时保持强鲁棒性
- 适合关注模型安全性与实用性的研究者
本文提出显著信息保持的对抗训练(SIP-AT),旨在缓解传统对抗训练带来的鲁棒性-准确性权衡问题。SIP-AT利用图像显著区域指导对抗训练过程,确保标注者认为有意义的脆弱特征在训练中不受扰动,使模型能在不牺牲整体鲁棒性的前提下学习高度预测性的非鲁棒特征。该方法兼容人工与自动生成的显著性估计,可无缝融入人工驱动的模型开发流程,无需依赖额外人工数据。我们在多个数据集和网络架构上进行了实验,结果表明SIP-AT在不同epsilon水平攻击下均能提升模型的干净准确率并维持高鲁棒性。我们进一步开展观察实验,测量人类识别扰动图像的速率,帮助直观理解对抗攻击强度,并凸显低epsilon鲁棒性的关键作用。结果验证了SIP-AT的有效性,并揭示了不同强度对抗样本带来的风险。
原文摘要 · Abstract (English)
In this work we introduce Salient Information Preserving Adversarial Training (SIP-AT), an intuitive method for relieving the robustness-accuracy trade-off incurred by traditional adversarial training. SIP-AT uses salient image regions to guide the adversarial training process in such a way that fragile features deemed meaningful by an annotator remain unperturbed during training, allowing models to learn highly predictive non-robust features without sacrificing overall robustness. This technique is compatible with both human-based and automatically generated salience estimates, allowing SIP-AT to be used as a part of human-driven model development without forcing SIP-AT to be reliant upon additional human data. We perform experiments across multiple datasets and architectures and demonstrate that SIP-AT is able to boost the clean accuracy of models while maintaining a high degree of robustness against attacks at multiple epsilon levels. We complement our central experiments with an observational study measuring the rate at which human subjects successfully identify perturbed images. This study helps build a more intuitive understanding of adversarial attack strength and demonstrates the heightened importance of low-epsilon robustness. Our results demonstrate the efficacy of SIP-AT and provide valuable insight into the risks posed by adversarial samples of various strengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。