arXiv:2410.12076cs.LGcs.CR2024-10被引 2

批判性审视对抗攻击的假设,揭示防御研究停滞的根源。

Taking off the Rose-Tinted Glasses: A Critical Look at Adversarial ML Through the Lens of Evasion Attacks

  • 从系统安全视角重构对抗攻击评估框架
  • 指出现有攻击假设脱离真实场景,导致防御要求不切实际
  • 适合关注AI安全落地的开发者与研究者

过去十年间,机器学习模型在对抗环境中的脆弱性受到学术界广泛关注,催生了大量攻击与防御方法。然而,尽管攻击手段不断拓展,防御研究却陷入停滞。当前,我们仍未找到超越额外训练之外的有效防护方案。随着生成式AI与大语言模型的快速发展,这一问题日益凸显。本文认为,攻击方过于宽松、防御方过于严格的威胁建模,严重制约了防御技术的发展。通过分析神经网络的对抗规避攻击,我们质疑了‘攻击可绕过任何未显式构建的防御’这一普遍假设。这些被论文接受机制所认可的假设,实际上使攻击远离真实场景。相应地,新防御必须近乎完美且嵌入模型本身才能通过测试。但现实中,机器学习模型只是更大系统的一部分。本文从系统安全角度重新审视对抗机器学习,探讨其对新兴AI范式的影响。

原文摘要 · Abstract (English)

The vulnerability of machine learning models in adversarial scenarios has garnered significant interest in the academic community over the past decade, resulting in a myriad of attacks and defenses. However, while the community appears to be overtly successful in devising new attacks across new contexts, the development of defenses has stalled. After a decade of research, we appear no closer to securing AI applications beyond additional training. Despite a lack of effective mitigations, AI development and its incorporation into existing systems charge full speed ahead with the rise of generative AI and large language models. Will our ineffectiveness in developing solutions to adversarial threats further extend to these new technologies? In this paper, we argue that overly permissive attack and overly restrictive defensive threat models have hampered defense development in the ML domain. Through the lens of adversarial evasion attacks against neural networks, we critically examine common attack assumptions, such as the ability to bypass any defense not explicitly built into the model. We argue that these flawed assumptions, seen as reasonable by the community based on paper acceptance, have encouraged the development of adversarial attacks that map poorly to real-world scenarios. In turn, new defenses evaluated against these very attacks are inadvertently required to be almost perfect and incorporated as part of the model. But do they need to? In practice, machine learning models are deployed as a small component of a larger system. We analyze adversarial machine learning from a system security perspective rather than an AI perspective and its implications for emerging AI paradigms.

对抗攻击系统安全模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。