针对大模型越狱检测评估难题,提出自适应多维度框架,精准识别不同场景下的越狱行为。
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
- 根据场景动态调整评估维度,突破传统统一标准的局限性。
- 在14个场景数据集上达F1 0.917,较现有方法提升6%。
- 适合安全测试、越狱研究者使用,尤其关注真实场景适配性。
准确的越狱评估对大模型红队测试和越狱研究至关重要。主流方法依赖二分类(字符串匹配、毒性文本分类器、基于LLM的方法),仅输出“是/否”标签,无法量化危害严重程度。新兴的多维评估框架(如安全违规、相对真实性与信息量)虽采用统一标准,但在不同场景中存在不适用问题(如“相对真实性”对“仇恨言论”场景无意义),影响评估准确性。为此,我们提出SceneJailEval:(1) 首个场景自适应的多维度越狱评估框架,克服现有方法“一刀切”的缺陷,具备强可扩展性,可无缝适配定制或新兴场景;(2) 构建包含14个场景的新型数据集,涵盖丰富越狱变体与地区案例,填补高质量、综合性基准的长期空白;(3) 在全场景数据集上实现F1 0.917(较SOTA提升6%),在JBB数据集上达F1 0.995(提升3%),突破异构场景下评估精度瓶颈,确立领先优势。
原文摘要 · Abstract (English)
Accurate jailbreak evaluation is critical for LLM red team testing and jailbreak research. Mainstream methods rely on binary classification (string matching, toxic text classifiers, and LLM-based methods), outputting only "yes/no" labels without quantifying harm severity. Emerged multi-dimensional frameworks (e.g., Security Violation, Relative Truthfulness and Informativeness) use unified evaluation standards across scenarios, leading to scenario-specific mismatches (e.g., "Relative Truthfulness" is irrelevant to "hate speech"), undermining evaluation accuracy. To address these, we propose SceneJailEval, with key contributions: (1) A pioneering scenario-adaptive multi-dimensional framework for jailbreak evaluation, overcoming the critical "one-size-fits-all" limitation of existing multi-dimensional methods, and boasting robust extensibility to seamlessly adapt to customized or emerging scenarios. (2) A novel 14-scenario dataset featuring rich jailbreak variants and regional cases, addressing the long-standing gap in high-quality, comprehensive benchmarks for scenario-adaptive evaluation. (3) SceneJailEval delivers state-of-the-art performance with an F1 score of 0.917 on our full-scenario dataset (+6% over SOTA) and 0.995 on JBB (+3% over SOTA), breaking through the accuracy bottleneck of existing evaluation methods in heterogeneous scenarios and solidifying its superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。