arXiv:2507.05538cs.AIcs.CR2025-07被引 9

AI红队测试需从单点漏洞检测升级为系统级风险评估

Red Teaming AI Red Teaming

  • 提出宏观系统与微观模型双层红队框架
  • 强调需关注技术与社会因素的交互影响
  • 适合关注AI治理与系统安全的研究者

红队测试已从军事应用演变为网络安全和人工智能领域的通用方法。本文批判性审视了当前AI红队测试实践:尽管在人工智能治理中广受欢迎,但其实际执行已偏离原始目标——即作为批判性思维训练,转而聚焦生成式AI中的模型级缺陷。现有红队工作主要关注单一模型漏洞,忽视了模型、用户与环境复杂互动所引发的社技术系统及涌现行为。为弥补这一不足,我们提出一个双层级红队框架:涵盖整个AI开发周期的宏观系统红队测试,以及针对具体模型的微观红队测试。基于网络安全经验与系统理论,进一步提出六项建议,强调有效红队测试需由多职能团队执行,系统性识别新兴风险、体系性漏洞及技术与社会因素的相互作用。

原文摘要 · Abstract (English)

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming. Drawing on cybersecurity experience and systems theory, we further propose a set of six recommendations. In these, we emphasize that effective AI red teaming requires multifunctional teams that examine emergent risks, systemic vulnerabilities, and the interplay between technical and social factors.

AI治理红队测试系统安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。