arXiv:2506.05376cs.CRcs.AI2025-06被引 4

聚焦系统级安全的红队测试路线图,提升真实场景下的模型防护能力。

A Red Teaming Roadmap Towards System-Level Safety

  • 以产品安全规范为优先,而非抽象伦理问题。
  • 模拟真实攻击者行为,构建更具代表性的威胁模型。
  • 将模型部署后的系统级防护(如用户检测)纳入红队测试范畴。

大型语言模型(LLM)的安全防护机制,如请求拒绝功能,已成为防止滥用的主流策略。在对抗机器学习与AI安全的交叉领域,针对拒绝训练的先进LLM进行红队测试已有效识别出关键漏洞。然而,我们认为当前大量关于LLM红队测试的会议投稿并未聚焦正确的研究问题。首先,应优先测试是否符合明确的产品安全规范,而非抽象的社会偏见或伦理原则;其次,红队测试应优先考虑反映不断扩展的风险格局和真实攻击者行为的现实威胁模型;最后,我们主张系统级安全是推动红队测试研究前进的必要步骤,因为模型在部署环境中既带来新威胁,也提供新的缓解手段(例如检测并封禁恶意用户)。采纳这些优先事项,才能有效应对快速发展的AI所带来及未来将出现的新威胁。

原文摘要 · Abstract (English)

Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.

红队测试系统安全LLM防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。