arXiv:2506.00781cs.AI2025-06NeurIPS被引 12

用原则组合自动化生成越狱提示,大幅提升模型安全漏洞发现能力

CoP: Agentic Red-teaming for Large Language Models using Composition of Principles

  • 通过人类提供的越狱原则,由智能体自动组合生成攻击策略
  • 在主流大模型上将越狱成功率提升至此前最高水平的19.0倍
  • 适合安全研究人员和模型开发者用于提前发现潜在风险

大型语言模型(LLMs)的快速发展推动了多个领域的应用,从开源到专有模型均有涉及。然而,越狱攻击——即通过诱导目标模型输出有害或高风险内容来突破安全对齐与用户合规性——正日益成为紧迫问题。红队测试旨在发布前沿AI技术前主动探测潜在风险和易出错场景。本文提出一种基于原则组合(CoP)框架的智能体工作流,使用户通过提供一组红队测试原则,由AI代理自动编排有效红队策略并生成越狱提示。与现有方法不同,CoP框架提供统一且可扩展的结构,能够整合与调度人类提供的红队原则,实现新红队策略的自动化发现。在主流大模型上的测试表明,CoP揭示了前所未有的安全风险,发现了新的越狱提示,并将最佳单轮攻击成功率提升至此前最高水平的19.0倍。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have spurred transformative applications in various domains, ranging from open-source to proprietary LLMs. However, jailbreak attacks, which aim to break safety alignment and user compliance by tricking the target LLMs into answering harmful and risky responses, are becoming an urgent concern. The practice of red-teaming for LLMs is to proactively explore potential risks and error-prone instances before the release of frontier AI technology. This paper proposes an agentic workflow to automate and scale the red-teaming process of LLMs through the Composition-of-Principles (CoP) framework, where human users provide a set of red-teaming principles as instructions to an AI agent to automatically orchestrate effective red-teaming strategies and generate jailbreak prompts. Distinct from existing red-teaming methods, our CoP framework provides a unified and extensible framework to encompass and orchestrate human-provided red-teaming principles to enable the automated discovery of new red-teaming strategies. When tested against leading LLMs, CoP reveals unprecedented safety risks by finding novel jailbreak prompts and improving the best-known single-turn attack success rate by up to 19.0 times.

红队测试越狱攻击智能体安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。