arXiv:2501.18837cs.CLcs.AI2025-01被引 192

用规则约束训练分类器,有效抵御大规模越狱攻击

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

  • 用自然语言规则生成合成数据训练安全分类器
  • 3000小时红队测试中未发现有效越狱攻击
  • 兼顾安全性与部署效率,适合生产环境使用

大型语言模型易受通用越狱攻击影响——即通过系统性提示绕过模型安全机制,执行需大量交互的有害行为(如大规模制造违禁品)。为此,我们提出宪法分类器:基于自然语言规则(即“宪法”)生成合成数据训练的安全防护机制。在超过3000小时的红队测试中,几乎无法找到能从早期受保护模型提取信息、达到未受保护模型水平的通用越狱策略。自动化评估显示,增强版分类器对未见领域的特定越狱攻击具有强防御能力。该方法兼具实用性,仅带来0.38%的生产流量拒绝率增加和23.7%的推理开销,证明在保障安全的同时实现可部署性是可行的。

原文摘要 · Abstract (English)

Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable.

安全防护越狱攻击分类器部署效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。