arXiv:2409.15398cs.CRcs.AI2024-09被引 22

构建实用框架,系统分析生成式AI单轮攻击漏洞

Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI

  • 提出攻击图谱框架,聚焦单轮输入攻击的实战分析
  • 梳理防御开发中的关键挑战与未解问题
  • 专为一线安全实践者设计,填补学术与落地鸿沟

随着生成式AI特别是大语言模型(LLMs)在生产环境中广泛应用,自然语言和多模态系统中涌现出新的攻击面和漏洞,引发对对抗性威胁的关注。红队测试在主动发现系统弱点方面日益重要,而蓝队则致力于抵御此类攻击。尽管学术界对生成式AI的对抗风险兴趣日增,但针对实际应用场景的评估与缓解指导仍十分有限。为此,本文贡献包括:(1) 对生成式AI安全防护中红/蓝队策略的实践性分析;(2) 识别防御开发与评估中的关键挑战与开放问题;(3) 构建攻击图谱(Attack Atlas),提供一种直观的单轮输入攻击分析方法,推动其在实践中应用。本工作旨在弥合学术洞见与实际安全措施之间的差距。

原文摘要 · Abstract (English)

As generative AI, particularly large language models (LLMs), become increasingly integrated into production applications, new attack surfaces and vulnerabilities emerge and put a focus on adversarial threats in natural language and multi-modal systems. Red-teaming has gained importance in proactively identifying weaknesses in these systems, while blue-teaming works to protect against such adversarial attacks. Despite growing academic interest in adversarial risks for generative AI, there is limited guidance tailored for practitioners to assess and mitigate these challenges in real-world environments. To address this, our contributions include: (1) a practical examination of red- and blue-teaming strategies for securing generative AI, (2) identification of key challenges and open questions in defense development and evaluation, and (3) the Attack Atlas, an intuitive framework that brings a practical approach to analyzing single-turn input attacks, placing it at the forefront for practitioners. This work aims to bridge the gap between academic insights and practical security measures for the protection of generative AI systems.

生成式AI安全红队测试对抗攻击实战框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。