提出真实数据重叠下的模型删忆评估基准,解决误删公共知识的问题
DUSK: Do Not Unlearn Shared Knowledge
- 构建含共享事实的文本集,模拟新闻与维基百科内容重叠场景
- 9种现有删忆方法均无法精准删除特定文本而不伤及共用知识
- 适合关注模型隐私保护与知识保留平衡的研究者使用
大型语言模型在实际应用中引发对版权或敏感数据被滥用的担忧。机器删忆旨在移除‘遗忘’数据的同时保留‘保留’数据的效用和信息。然而,现有评估通常假设遗忘与保留数据完全不重叠,忽略了现实场景中的内容重叠。例如,一篇关于日本地震的新闻需被删忆,但该事件在维基百科上也存在客观描述。理想删忆应仅移除新闻的特定表述,保留公众认可的事实。本文提出DUSK基准,用于评估在真实数据重叠情况下的删忆方法。DUSK构造了以不同风格描述相同事实的文档集,部分信息跨集共享,部分内容唯一。当指定一个集合进行删忆时,理想方法应移除其独有内容,同时保留共享事实。我们定义七项评估指标来检验删忆方法是否实现选择性移除。对九种近期删忆方法的评估发现:多数方法虽能移除表层文本,却难以消除深层上下文知识而不损害共享内容。我们公开发布DUSK,以支持更精确可靠的删忆技术发展。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in real-world applications, raising concerns about the unauthorized use of copyrighted or sensitive data. Machine unlearning aims to remove such 'forget' data while preserving utility and information from the 'retain' set. However, existing evaluations typically assume that forget and retain sets are fully disjoint, overlooking realistic scenarios where they share overlapping content. For instance, a news article may need to be unlearned, even though the same event, such as an earthquake in Japan, is also described factually on Wikipedia. Effective unlearning should remove the specific phrasing of the news article while preserving publicly supported facts. In this paper, we introduce DUSK, a benchmark designed to evaluate unlearning methods under realistic data overlap. DUSK constructs document sets that describe the same factual content in different styles, with some shared information appearing across all sets and other content remaining unique to each. When one set is designated for unlearning, an ideal method should remove its unique content while preserving shared facts. We define seven evaluation metrics to assess whether unlearning methods can achieve this selective removal. Our evaluation of nine recent unlearning methods reveals a key limitation: while most can remove surface-level text, they often fail to erase deeper, context-specific knowledge without damaging shared content. We release DUSK as a public benchmark to support the development of more precise and reliable unlearning techniques for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。