通过语义分析与流畅度检测,识别优化型越狱攻击
SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

- 结合跨层分布距离与困惑度的混合流畅度测量
- 利用梯度匹配分析有害语义,识别伪装恶意提示
- 对多种优化型越狱攻击检测准确率显著提升
尽管大量研究致力于使大语言模型(LLMs)与人类价值观对齐并确保安全部署,但近期工作揭示,LLMs仍易受对抗性越狱攻击影响,可绕过安全防护生成有害内容。现有防御方法在应对广泛优化型越狱机制时效果有限,这些机制能生成高度流畅或语义混淆的恶意提示。为此,我们提出统一检测框架SAFEGuard,结合基于跨层分布距离和困惑度的混合流畅度测量,以及通过梯度匹配分析有害语义。该方法基于关键观察:高流畅度提示在恶意意图上接近有害提示,而语义混淆提示常引入无意义标记序列。评估表明,SAFEGuard在不同优化型越狱攻击中均显著优于现有基线,检测准确率明显提升,证明其对演化中越狱攻击的有效性。
原文摘要 · Abstract (English)
Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to detect jailbreaks but they are limited in their effectiveness to counter wide-range optimization-based jailbreak mechanisms that can yield highly fluency-optimized or harmful semantic obfuscated prompts. To tackle this challenge, we propose a unified detection framework SAFEGuard which incorporates a hybrid fluency measurement based on cross-layer distribution distance and perplexity, and the analysis of harmful semantics through gradient matching. Our method is grounded in a paramount observation: high fluency prompts maintain their malicious intention close to harmful prompts while harmful semantic obfuscated prompts often inject gibberish token sequences. Our evaluation demonstrates that SAFEGuard consistently outperforms state-of-the-art baselines and achieves significant improvement in accuracy across different optimization-based jailbreaks. This underscores the effectiveness of SAFEGuard against evolving jailbreak attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。