用攻击样本生成可更新的安全测试集,提升评测公平性与可持续性。
Jailbreak Distillation: Renewable Safety Benchmarking
- 从攻击方法中提炼高质量提示,自动构建安全评测集。
- 新基准在13个模型上表现优于现有方法,通用性强且多样性高。
- 只需少量人力即可重跑生成新版本,适合长期安全评估。
大型语言模型在关键应用中快速部署,亟需可靠的安全部署评测。我们提出一种名为Jailbreak Distillation(JBDistill)的新框架,将越狱攻击“蒸馏”为高质量、易更新的安全评测基准。该框架利用少量开发模型和现有越狱攻击算法生成候选提示池,再通过提示选择算法筛选出有效子集作为评测基准。相比现有方法,本方案确保不同模型间使用一致的评测提示,提升比较公平性与可复现性;仅需极少人工干预即可重运行管道并生成更新版基准,缓解评测饱和与污染问题。大量实验表明,该基准在13个未参与构建的多样化模型上均表现稳健,涵盖专有模型、专用模型及新一代模型,显著优于现有安全评测基准,同时保持高区分度与多样性。该框架为安全评测提供了高效、可持续且可适应的解决方案。
原文摘要 · Abstract (English)
Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a novel benchmark construction framework that "distills" jailbreak attacks into high-quality and easily-updatable safety benchmarks. JBDistill utilizes a small set of development models and existing jailbreak attack algorithms to create a candidate prompt pool, then employs prompt selection algorithms to identify an effective subset of prompts as safety benchmarks. JBDistill addresses challenges in existing safety evaluation: the use of consistent evaluation prompts across models ensures fair comparisons and reproducibility. It requires minimal human effort to rerun the JBDistill pipeline and produce updated benchmarks, alleviating concerns on saturation and contamination. Extensive experiments demonstrate our benchmarks generalize robustly to 13 diverse evaluation models held out from benchmark construction, including proprietary, specialized, and newer-generation LLMs, significantly outperforming existing safety benchmarks in effectiveness while maintaining high separability and diversity. Our framework thus provides an effective, sustainable, and adaptable solution for streamlining safety evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。