arXiv:2410.15042cs.LGcs.AI2024-10综述被引 22

系统梳理对抗训练方法,帮模型抵御细微扰动攻击

Adversarial Training: A Survey

  • 将对抗样本融入训练过程提升模型鲁棒性
  • 从数据增强、网络设计、训练配置三方面总结技术进展
  • 适合关注模型安全与防御的科研人员参考

对抗训练(AT)通过在训练中引入对抗样本——即添加难以察觉的扰动但能显著影响模型预测的输入——来提升深度神经网络对各类对抗攻击的鲁棒性。尽管近期研究已证实其有效性,但相关进展尚缺乏全面综述。本文系统回顾了近年来代表性研究成果,首先介绍AT的实现流程与实际应用,随后从数据增强、网络结构设计、训练配置三个维度展开全面分析,并讨论当前面临的共性挑战,最后提出若干有前景的未来研究方向。

原文摘要 · Abstract (English)

Adversarial training (AT) refers to integrating adversarial examples -- inputs altered with imperceptible perturbations that can significantly impact model predictions -- into the training process. Recent studies have demonstrated the effectiveness of AT in improving the robustness of deep neural networks against diverse adversarial attacks. However, a comprehensive overview of these developments is still missing. This survey addresses this gap by reviewing a broad range of recent and representative studies. Specifically, we first describe the implementation procedures and practical applications of AT, followed by a comprehensive review of AT techniques from three perspectives: data enhancement, network design, and training configurations. Lastly, we discuss common challenges in AT and propose several promising directions for future research.

对抗训练模型鲁棒性深度学习安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。