构建首个可解释的越狱攻击评估基准,支持多语言与复杂场景测试
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- 提出多智能体推理框架,实现1-10分细粒度可解释评分
- 涵盖3.5万+指令微调数据与4.5千+风险样本,覆盖十种语言
- 适用于安全模型评估、攻防研究,尤其适合零样本场景
尽管大模型安全防护能力持续提升,但评估仍面临可解释性不足与泛化性差的问题,现有方法常仅作直接判断、在复杂场景中表现不佳(如GPT-4 F1得分低、多语言存在偏见)。为此,我们提出JAILJUDGE,一个综合性基准,包含合成、对抗、真实世界及多语言等多样化风险场景,配备高质量人工标注数据集。数据集包含超过35,000条指令微调数据(含推理解释)和4,500+标注样本用于风险评估,以及跨10种语言的6,000+多语言样本。为增强评估可解释性,我们设计了JailJudge多智能体框架,实现1-10分细粒度评分。该框架支持构建指令微调的真值,并推动开发出端到端的判别模型JAILJUDGE Guard,提供推理过程且无需支付API费用。此外,我们引入攻击增强器JailBoost和防御系统GuardShield,均基于JAILJUDGE Guard。实验表明,JailJudge方法在多种模型(如GPT-4、Llama-Guard)及零样本场景下均达到领先性能。JailBoost使攻击效果提升29.24%,GuardShield将防御误报率从40.46%降至0.15%。
原文摘要 · Abstract (English)
Despite advancements in enhancing LLM safety against jailbreak attacks, evaluating LLM defenses remains a challenge, with current methods often lacking explainability and generalization to complex scenarios, leading to incomplete assessments (e.g., direct judgment without reasoning, low F1 score of GPT-4 in complex cases, bias in multilingual scenarios). To address this, we present JAILJUDGE, a comprehensive benchmark featuring diverse risk scenarios, including synthetic, adversarial, in-the-wild, and multilingual prompts, along with high-quality human-annotated datasets. The JAILJUDGE dataset includes over 35k+ instruction-tune data with reasoning explainability and JAILJUDGETEST, a 4.5k+ labeled set for risk scenarios, and a 6k+ multilingual set across ten languages. To enhance evaluation with explicit reasoning, we propose the JailJudge MultiAgent framework, which enables explainable, fine-grained scoring (1 to 10). This framework supports the construction of instruction-tuning ground truth and facilitates the development of JAILJUDGE Guard, an end-to-end judge model that provides reasoning and eliminates API costs. Additionally, we introduce JailBoost, an attacker-agnostic attack enhancer, and GuardShield, a moderation defense, both leveraging JAILJUDGE Guard. Our experiments demonstrate the state-of-the-art performance of JailJudge methods (JailJudge MultiAgent, JAILJUDGE Guard) across diverse models (e.g., GPT-4, Llama-Guard) and zero-shot scenarios. JailBoost and GuardShield significantly improve jailbreak attack and defense tasks under zero-shot settings, with JailBoost enhancing performance by 29.24% and GuardShield reducing defense ASR from 40.46% to 0.15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。