arXiv:2505.18556cs.CLcs.AI2025-05EMNLP被引 9

通过意图操控破解大模型内容安全防线,暴露其防护漏洞。

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

  • 设计两阶段提示优化框架,将有害提问转化为隐性指令
  • 在多个模型上实现86.75%~97.12%的越狱成功率
  • 揭示意图感知防护机制的可被绕过风险,适合安全研究者参考

意图检测是自然语言理解的核心组件,已成为保护大语言模型(LLMs)的重要机制。尽管已有研究利用意图检测提升模型的安全防护,有效抵御内容级越狱攻击,但这些基于意图的防护机制在恶意操纵下的鲁棒性仍缺乏深入探索。本文研究了意图感知防护系统的脆弱性,发现大模型具备隐式的意图识别能力。为此,我们提出两阶段的意图驱动提示优化框架IntentPrompt:首先将有害请求转化为结构化提纲,再通过反馈循环迭代优化提示,将其重构为陈述式叙事,以提升红队测试中的越狱成功率。在四个公开基准和多种黑盒大模型上的实验表明,该框架持续优于多个前沿越狱方法,并能绕过先进的意图分析(IA)与思维链(CoT)防御机制。特别地,'FSTR+SPIN'变体在o1模型上对CoT防御的攻击成功率达88.25%~96.54%,在GPT-4o模型上对IA防御的成功率为86.75%~97.12%。这些结果凸显大模型安全机制的关键缺陷,表明意图操控正构成内容监管的重大挑战。

原文摘要 · Abstract (English)

Intent detection, a core component of natural language understanding, has considerably evolved as a crucial mechanism in safeguarding large language models (LLMs). While prior work has applied intent detection to enhance LLMs' moderation guardrails, showing a significant success against content-level jailbreaks, the robustness of these intent-aware guardrails under malicious manipulations remains under-explored. In this work, we investigate the vulnerability of intent-aware guardrails and demonstrate that LLMs exhibit implicit intent detection capabilities. We propose a two-stage intent-based prompt-refinement framework, IntentPrompt, that first transforms harmful inquiries into structured outlines and further reframes them into declarative-style narratives by iteratively optimizing prompts via feedback loops to enhance jailbreak success for red-teaming purposes. Extensive experiments across four public benchmarks and various black-box LLMs indicate that our framework consistently outperforms several cutting-edge jailbreak methods and evades even advanced Intent Analysis (IA) and Chain-of-Thought (CoT)-based defenses. Specifically, our "FSTR+SPIN" variant achieves attack success rates ranging from 88.25% to 96.54% against CoT-based defenses on the o1 model, and from 86.75% to 97.12% on the GPT-4o model under IA-based defenses. These findings highlight a critical weakness in LLMs' safety mechanisms and suggest that intent manipulation poses a growing challenge to content moderation guardrails.

安全漏洞越狱攻击意图检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。