提出更严格的测试框架,检验大模型删数据后是否仍可被强行还原。
Stress Testing Unlearning Algorithms

- 设计新测试流程,主动尝试从模型中提取已删除的数据
- 新增边界问题评估,检测与删除内容相近的查询表现
- 适合关注模型安全性和数据隐私的研究者使用
近期,机器遗忘(machine unlearning)——即消除特定训练数据对模型的影响——受到越来越多关注。在大语言模型(LLMs)中,由于输入输出的模糊性,遗忘尤为困难。因此,严格评估对于保障安全性和可用性、推动方法进步至关重要。我们发现现有遗忘基准存在两大缺陷:(1) 未主动测试被遗忘信息是否仍可被强制提取;(2) 未能评估边界问题上的性能保留,即语义上接近被遗忘内容的良性查询。为此,我们提出 WMDP++,作为 WMDP 的扩展,通过引入针对遗忘信息的定向提取和系统化边界问题评估,填补上述空白。WMDP++ 为评估大语言模型中的遗忘效果提供了更严格、更丰富的基准。
原文摘要 · Abstract (English)
Recently, machine unlearning, the removal of specific training data influence from a model, has gained increasing attention. In large language models (LLMs), unlearning is particularly challenging due to the ambiguity of inputs and outputs. Con- sequently, rigorous evaluation is critical for assessing both safety and utility, and for driving progress in unlearning meth- ods. We identify two key shortcomings in existing unlearning benchmarks: (1) they do not actively test whether unlearned information can still be forcibly extracted, and (2) they fail to evaluate performance preservation on boundary questions, be- nign queries that are semantically close to the unlearned con- tent. Here we introduce WMDP++, an extension of WMDP that addresses these gaps by incorporating targeted extrac- tion of unlearned information and systematic evaluation on boundary questions. WMDP++ provides a more stringent and informative benchmark for evaluating unlearning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。