提出可信赖的越狱攻击评估框架,提升成功率测量准确性。
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

- 基于信息瓶颈理论,用双反馈优化过滤干扰内容
- 在330个越狱样本上准确率达97.27%,显著优于现有方法
- 推出轻量版模型,支持大规模高效评估
当前大语言模型越狱攻击的评估缺乏一致标准,导致成功率估计不可靠。我们提出JailMeter,一种基于证据的评估框架,更真实地衡量越狱效果。受信息瓶颈理论启发,该框架采用双反馈优化,从模型回复中过滤越狱噪声,同时保留与原始恶意问题相关的语义内容。该过程生成简洁证据,仅当回复完整捕捉恶意意图并给出完整答案时才判定为成功越狱,体现对安全对齐的有效绕过。我们在包含330个人工标注、未被拒绝的越狱实例的基准JailMeter-Eva上进行评估,JailMeter达到97.27%的准确率,显著优于现有方法。为进一步支持大规模评估,我们还将JailMeter蒸馏为小型语言模型JailMeter_SLM,保持相近可靠性的同时大幅降低计算开销。代码与数据集已开源。
原文摘要 · Abstract (English)
The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness. Inspired by the Information Bottleneck theory, JailMeter applies dual-feedback optimization to filter jailbreak noise from model responses while preserving content relevant to the original malicious question. This process produces concise evidence for a rigorous assessment under which an attack is validated only when the response captures the malicious intent and delivers a complete answer, thereby signaling a substantive bypass of model safety alignment. We evaluate JailMeter on JailMeter-Eva, a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. JailMeter achieves an accuracy of 97.27%, substantially outperforming existing evaluation methods. To support large-scale evaluation, we further distill JailMeter into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs. Code and dataset are available at https://github.com/Magi2B0y/JailMeter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。