arXiv:2504.04976cs.CL2025-04被引 1

从训练数据域角度重新分类大模型越狱漏洞,揭示其内在弱点。

A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models

  • 按模型训练域划分越狱攻击类型,突破传统提示构造分类。
  • 提出四类漏洞:泛化错配、目标冲突、鲁棒性缺失与混合攻击。
  • 帮助安全研究者理解模型缺陷本质,适合模型安全方向读者。

大语言模型(LLMs)在开放世界机器学习中占据核心地位。尽管具备出色的自然语言处理能力,但面临一致性问题、幻觉及越狱漏洞等挑战。越狱指通过精心设计的提示绕过对齐防护,生成不安全输出,破坏模型完整性。本文聚焦越狱漏洞,提出基于模型训练域的新分类体系。通过泛化、目标与鲁棒性缺口刻画对齐失败。主要贡献是将越狱视角置于训练中出现的语言域之上,揭示现有方法局限,并据此分类越狱攻击所利用的模型缺陷。不同于传统按提示构造方式(如提示模板)分类,本方法深化对模型行为的理解。提出包含四类的分类体系:泛化错配、目标冲突、对抗鲁棒性缺失与混合攻击,揭示越狱漏洞的本质。最后总结关键经验教训。

原文摘要 · Abstract (English)

The study of large language models (LLMs) is a key area in open-world machine learning. Although LLMs demonstrate remarkable natural language processing capabilities, they also face several challenges, including consistency issues, hallucinations, and jailbreak vulnerabilities. Jailbreaking refers to the crafting of prompts that bypass alignment safeguards, leading to unsafe outputs that compromise the integrity of LLMs. This work specifically focuses on the challenge of jailbreak vulnerabilities and introduces a novel taxonomy of jailbreak attacks grounded in the training domains of LLMs. It characterizes alignment failures through generalization, objectives, and robustness gaps. Our primary contribution is a perspective on jailbreak, framed through the different linguistic domains that emerge during LLM training and alignment. This viewpoint highlights the limitations of existing approaches and enables us to classify jailbreak attacks on the basis of the underlying model deficiencies they exploit. Unlike conventional classifications that categorize attacks based on prompt construction methods (e.g., prompt templating), our approach provides a deeper understanding of LLM behavior. We introduce a taxonomy with four categories -- mismatched generalization, competing objectives, adversarial robustness, and mixed attacks -- offering insights into the fundamental nature of jailbreak vulnerabilities. Finally, we present key lessons derived from this taxonomic study.

模型安全越狱漏洞分类体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。