arXiv:2512.24044cs.CRcs.AI2025-12ACL被引 3

首次系统评估越狱攻击在完整部署链中的成功率,发现安全过滤器可有效拦截多数攻击。

Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?

  • 测试越狱攻击在输入输出双阶段过滤下的表现,覆盖完整推理流程
  • 近全数越狱手法被至少一种过滤器识别,表明过去评估高估了攻击成功率
  • 需优化检测精度与召回率平衡,兼顾安全与用户体验

随着大语言模型(LLMs)的广泛应用,确保其安全使用至关重要。越狱攻击——即通过对抗性提示绕过模型对齐机制以触发有害输出——构成重大风险,现有研究报道其在常见LLMs上具有较高成功率。然而,以往评估仅关注模型本身,忽略了实际部署中通常包含的内容安全过滤机制。为此,我们首次系统评估针对LLM安全对齐的越狱攻击,考察其在完整推理管道中的成功率,涵盖输入与输出过滤阶段。研究发现:第一,几乎所有测试的越狱技术均可被至少一种安全过滤器检测到,表明先前评估可能过高估计了这些攻击的实际成功概率;第二,尽管安全过滤器具备检测能力,但在召回率与精确率之间仍需更好权衡,以进一步提升防护效果与用户使用体验。本文揭示关键缺口,呼吁进一步优化大语言模型安全系统的检测准确性和可用性。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful outputs, present significant risks, with existing studies reporting high success rates in evading common LLMs. However, previous evaluations have focused solely on the models, neglecting the full deployment pipeline, which typically incorporates additional safety mechanisms like content moderation filters. To address this gap, we present the first systematic evaluation of jailbreak attacks targeting LLM safety alignment, assessing their success across the full inference pipeline, including both input and output filtering stages. Our findings yield two key insights: first, nearly all evaluated jailbreak techniques can be detected by at least one safety filter, suggesting that prior assessments may have overestimated the practical success of these attacks; second, while safety filters are effective in detection, there remains room to better balance recall and precision to further optimize protection and user experience. We highlight critical gaps and call for further refinement of detection accuracy and usability in LLM safety systems.

LLM安全越狱攻击内容过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。