arXiv:2510.13893cs.CLcs.AI2025-10被引 5

构建七类劫持策略分类体系,提升大模型安全防御能力

Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection

  • 提出分层分类体系,整合七类劫持技术机制
  • 实测显示不同攻击策略成功率差异显著,揭示模型漏洞特征
  • 构建首个意大利语多轮对抗对话数据集,支持渐进式攻击研究

劫持攻击对大语言模型的安全构成严重威胁。现有防御方法多聚焦单轮攻击,跨语言覆盖不足,且分类体系有限,难以全面捕捉攻击策略多样性或准确反映技术本质。为深入理解劫持技术的有效性,我们开展结构化红队测试。实验成果包括:第一,构建涵盖七类机制的分层分类体系——伪装、说服、权限提升、认知过载、混淆、目标冲突与数据污染,系统整合孤立研究并明确关联已有分类;第二,分析测试数据,揭示各类攻击的流行度与成功率,揭示其如何利用模型漏洞引发偏差;第三,以GPT-5作为评判器,验证基于分类引导提示在自动检测中的增益;第四,构建包含1364条多轮对抗对话的意大利语数据集,并基于新分类标注,支持渐进式恶意意图演化研究。

原文摘要 · Abstract (English)

Jailbreaking techniques pose a significant threat to the safety of Large Language Models (LLMs). Existing defenses typically focus on single-turn attacks, lack coverage across languages, and rely on limited taxonomies that either fail to capture the full diversity of attack strategies or emphasize risk categories rather than jailbreaking techniques. To advance the understanding of the effectiveness of jailbreaking techniques, we conducted a structured red-teaming challenge. The outcomes of our experiments are fourfold. First, we developed a comprehensive hierarchical taxonomy of jailbreak strategies that systematically consolidates techniques previously studied in isolation and harmonizes existing, partially overlapping classifications with explicit cross-references to prior categorizations. The taxonomy organizes jailbreak strategies into seven mechanism-oriented families: impersonation, persuasion, privilege escalation, cognitive overload, obfuscation, goal conflict, and data poisoning. Second, we analyzed the data collected from the challenge to examine the prevalence and success rates of different attack types, providing insights into how specific jailbreak strategies exploit model vulnerabilities and induce misalignment. Third, we benchmarked GPT-5 as a judge for jailbreak detection, evaluating the benefits of taxonomy-guided prompting for improving automatic detection. Finally, we compiled a new Italian dataset of 1364 multi-turn adversarial dialogues, annotated with our taxonomy, enabling the study of interactions where adversarial intent emerges gradually and succeeds in bypassing traditional safeguards.

大模型安全劫持检测分类体系多轮对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。