arXiv:2605.05058cs.CRcs.AI2026-05被引 4

系统梳理大模型越狱攻击与防御,提出多维评估框架

SoK: Robustness in Large Language Models against Jailbreak Attacks

论文配图:SoK: Robustness in Large Language Models against Jailbreak Attacks
图 1 · 摘自论文原文
  • 构建安全立方体框架,从多维度评估越狱攻防技术
  • 评测13种攻击和5种防御,揭示现有方法的优劣与局限
  • 适合关注大模型安全与可信性的研究者与工程师

大型语言模型(LLMs)虽取得显著进展,但仍易受越狱攻击影响,即通过对抗性提示诱导模型生成有害、不道德或违反政策的内容。此类攻击在高风险场景中威胁安全、信任与合规性。尽管已有多种攻防方法被提出,但现有评估手段不足,常依赖单一指标如攻击成功率,难以反映模型安全的多维特性。本文系统构建越狱攻击与防御的分类体系,提出安全立方体(Security Cube)统一多维评估框架。通过详细对比表梳理现有攻防技术,揭示关键洞见与开放挑战。基于该框架,对13种代表性攻击与5种防御进行基准测试,全面呈现越狱攻击、防御策略、自动化评测工具及模型漏洞现状。据此提炼核心发现,识别未解问题,并指明增强大模型抗越狱能力的未来方向。本研究旨在推动更鲁棒、可解释、可信的大语言模型系统发展。代码已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, in which adversarial prompts coerce models into generating harmful, unethical, or policy-violating outputs. Such attacks pose real-world risks, eroding safety, trust, and regulatory compliance in high-stakes applications. Although a variety of attack and defense methods have been proposed, existing evaluation practices are inadequate, often relying on narrow metrics like attack success rate that fail to capture the multidimensional nature of LLM security. In this paper, we present a systematic taxonomy of jailbreak attacks and defenses and introduce Security Cube, a unified, multi-dimensional framework for comprehensive evaluation of these techniques. We provide detailed comparison tables of existing attacks and defenses, highlighting key insights and open challenges across the literature. Leveraging Security Cube, we conduct benchmark studies on 13 representative attacks and 5 defenses, establishing a clear view of the current landscape encompassing jailbreak attacks, defenses, automated judges, and LLM vulnerabilities. Based on these evaluations, we distill critical findings, identify unresolved problems, and outline promising research directions for enhancing LLM robustness against jailbreak attacks. Our analysis aims to pave the way towards more robust, interpretable, and trustworthy LLM systems. Our code is available at Code.

大模型安全越狱攻击评估框架鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。