arXiv:2506.15699cs.LGcs.AI2025-06被引 11

构建更真实的遗忘评估基准,检验大模型删知识后是否真能防泄露。

BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap

  • 设计含重叠信息的遗忘-保留任务,模拟真实场景
  • 现有方法在新基准上性能普遍下降,部分简单方法反而更好
  • 适合关注模型安全与鲁棒性评估的研究者使用

机器遗忘有望通过事后移除敏感或有害信息来提升大语言模型的安全性。其核心挑战在于平衡遗忘质量(有效删除不良信息)与保留质量(保持其他通用任务性能)。然而,我们发现现有大模型遗忘基准中遗忘集与保留集高度分离,导致对遗忘方法效果的评估失真,易被良性扰动(如重学攻击)暴露本应已遗忘的知识。为此,我们提出$ exttt{BLUR}$:一个考虑遗忘-保留重叠的更真实大模型遗忘评估基准。该基准扩展了现有评测任务,引入混合查询和不同难度的重学数据集。尽管所考虑查询均为良性,但现有方法在$ exttt{BLUR}$上的表现显著下降,简单方法平均优于近期复杂方法。结果凸显了稳健评估的重要性,并指明未来研究方向。基准已公开于:https://huggingface.co/datasets/forgelab/BLUR

原文摘要 · Abstract (English)

Machine unlearning has the potential to improve the safety of large language models (LLMs) by removing sensitive or harmful information post hoc. A key challenge in unlearning involves balancing between forget quality (effectively unlearning undesirable information) and retain quality (maintaining good performance on other, general tasks). Unfortunately, as we show, current LLM unlearning benchmarks contain highly disparate forget and retain sets -- painting a false picture of the effectiveness of LLM unlearning methods. This can be particularly problematic because it opens the door for benign perturbations, such as relearning attacks, to easily reveal supposedly unlearned knowledge once models are deployed. To address this, we present $\texttt{BLUR}$: a benchmark for LLM unlearning that provides more realistic scenarios of forget-retain overlap. $\texttt{BLUR}$ significantly expands on existing unlearning benchmarks by providing extended evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on $\texttt{BLUR}$, with simple approaches performing better on average than more recent methods. These results highlight the importance of robust evaluation and suggest several important directions of future study. Our benchmark is publicly available at: https://huggingface.co/datasets/forgelab/BLUR

模型安全遗忘学习评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。