提出新评估框架,解决大模型越狱测试中的结果失真问题
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
- 构建针对性有害问题数据集与分步评估指南
- 使越狱成功率评估误差降低76%以上
- 适合安全研究者和模型评测团队使用
尽管越狱攻击作为构建安全大模型的有效红队测试手段日益受到关注,但现有评估系统设计缺陷导致其有效性评估存在显著偏差。基于2022年以来37项越狱研究的系统性测量,我们发现现有评估缺乏针对具体案例的标准,从而得出误导性结论。本文提出GuidedBench,包含精心筛选的有害问题数据集和集成详细案例评估指南的GuidedEval评估系统。实验表明,GuidedBench能更准确评估越狱表现,实现方法间可比性;GuidedEval将评估者间差异降低至少76.03%,确保评估可靠且可复现。我们揭示了现有基准失效的原因,并提出更优评估实践。
原文摘要 · Abstract (English)
Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness assessments. With a systematic measurement study based on 37 jailbreak studies since 2022, we find that existing evaluation systems lack case-specific criteria, resulting in misleading conclusions about their effectiveness and safety implications. In this paper, we introduce GuidedBench, a novel benchmark comprising a curated harmful question dataset and GuidedEval, an evaluation system integrated with detailed case-by-case evaluation guidelines. Experiments demonstrate that GuidedBench offers more accurate evaluations of jailbreak performance, enabling meaningful comparisons across methods. GuidedEval reduces inter-evaluator variance by at least 76.03%, ensuring reliable and reproducible evaluations. We reveal why existing jailbreak benchmarks fail to evaluate accurately and suggest better evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。