研究发现越狱输出的实用性能大幅下降,称为越狱税。
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
- 用良性题目构建可验证答案的越狱评估集
- 8种越狱方法在数学任务上准确率最高降92%
- 提出越狱税新指标,适合安全与评测研究者
越狱攻击旨在绕过大语言模型的安全限制以生成有害内容。本文探究现有越狱方法产生的输出是否真正有用。由于多数有害回答(如制爆指南)难以严格评估,我们通过将模型对生物、数学等良性易评话题拒答,构建具有已知真实答案的新越狱评估集。在五个实用性基准上对八种代表性越狱方法的评估显示,越狱后模型性能普遍下降,我们称之为“越狱税”。例如,所有测试越狱均能绕过对数学问题的拒答,但准确率最高下降92%。本工作提出越狱税作为人工智能安全的重要新指标,并公开了评估基准(https://github.com/ethz-spylab/jailbreak-tax)。
原文摘要 · Abstract (English)
Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually useful. For example, when jailbreaking a model to give instructions for building a bomb, does the jailbreak yield good instructions? Since the utility of most unsafe answers (e.g., bomb instructions) is hard to evaluate rigorously, we build new jailbreak evaluation sets with known ground truth answers, by aligning models to refuse questions related to benign and easy-to-evaluate topics (e.g., biology or math). Our evaluation of eight representative jailbreaks across five utility benchmarks reveals a consistent drop in model utility in jailbroken responses, which we term the jailbreak tax. For example, while all jailbreaks we tested bypass guardrails in models aligned to refuse to answer math, this comes at the expense of a drop of up to 92% in accuracy. Overall, our work proposes the jailbreak tax as a new important metric in AI safety, and introduces benchmarks to evaluate existing and future jailbreaks. We make the benchmark available at https://github.com/ethz-spylab/jailbreak-tax
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。