arXiv:2505.05064cs.LG2025-05

用文本水印评估大模型数据遗忘效果,更精准判断该删哪些内容。

WaterDrum: Watermarking for Data-centric Unlearning Metric

  • 基于文本水印设计数据导向的遗忘评估方法
  • 在语义相似数据场景下仍能准确衡量遗忘程度
  • 适合需要精确删除敏感或版权数据的场景

大语言模型的遗忘机制在实际应用中至关重要,需高效消除私密、版权或有害数据的影响。现有以模型性能为中心的遗忘评估指标,在遗忘集与保留集语义相近,或无法从头训练模型时,难以准确评估遗忘程度。本文提出首个面向大模型的数据中心式遗忘评估指标 WaterDrum,利用鲁棒文本水印克服上述局限。我们构建了包含不同数据相似度水平的新基准数据集,可用于通过 WaterDrum 严格评估遗忘算法。代码已开源于 https://github.com/lululu008/WaterDrum,新基准数据集发布于 https://huggingface.co/datasets/Glow-AI/WaterDrum-Ax。

原文摘要 · Abstract (English)

Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from some users. Existing utility-centric unlearning metrics (based on model utility) may fail to accurately evaluate the extent of unlearning in realistic settings such as when the forget and retain sets have semantically similar content and/or retraining the model from scratch on the retain set is impractical. This paper presents the first data-centric unlearning metric for LLMs called WaterDrum that exploits robust text watermarking to overcome these limitations. We introduce new benchmark datasets (with different levels of data similarity) for LLM unlearning that can be used to rigorously evaluate unlearning algorithms via WaterDrum. Our code is available at https://github.com/lululu008/WaterDrum and our new benchmark datasets are released at https://huggingface.co/datasets/Glow-AI/WaterDrum-Ax.

大模型数据遗忘水印技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。