现有大模型遗忘测试基准夸大效果,易被简单修改骗过
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
- 用轻微改动测试集,暴露遗忘后仍可获取敏感信息
- 原基准显示遗忘有效,实际模型性能却严重下降
- 适合关注隐私安全与评估可靠性的研究者参考
遗忘方法有望通过事后移除敏感或有害信息来提升大语言模型(LLM)的隐私与安全性。当前研究社区越来越依赖实证基准来评估这些方法的有效性。本文发现,现有基准对候选遗忘方法的效果呈现过于乐观且可能误导的评价。通过引入对多个流行基准的简单、无害修改,我们揭示出:本应被遗忘的信息仍可被访问,或遗忘过程对保留信息的性能损害远超原基准所反映的程度。我们指出,现有基准特别容易受到遗忘与保留信息间松散关联的影响。此外,基准中遗忘目标的模糊性也导致方法容易过拟合特定测试查询。基于此,我们呼吁社区谨慎解读基准结果作为进展的可靠指标,并提出多项未来研究建议。
原文摘要 · Abstract (English)
Unlearning methods have the potential to improve the privacy and safety of large language models (LLMs) by removing sensitive or harmful information post hoc. The LLM unlearning research community has increasingly turned toward empirical benchmarks to assess the effectiveness of such methods. In this paper, we find that existing benchmarks provide an overly optimistic and potentially misleading view on the effectiveness of candidate unlearning methods. By introducing simple, benign modifications to a number of popular benchmarks, we expose instances where supposedly unlearned information remains accessible, or where the unlearning process has degraded the model's performance on retained information to a much greater extent than indicated by the original benchmark. We identify that existing benchmarks are particularly vulnerable to modifications that introduce even loose dependencies between the forget and retain information. Further, we show that ambiguity in unlearning targets in existing benchmarks can easily lead to the design of methods that overfit to the given test queries. Based on our findings, we urge the community to be cautious when interpreting benchmark results as reliable measures of progress, and we provide several recommendations to guide future LLM unlearning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。