arXiv:2503.14900cs.CLcs.AI2025-03被引 5

让大模型遗忘特定训练数据,同时保持生成能力。

Deep Contrastive Unlearning for Language Models

  • 通过优化模型隐空间几何分布实现遗忘
  • 在真实数据集上显著优于基线方法
  • 适合需要隐私保护的模型更新场景

近年来大语言模型在理解文本和生成类人语言方面取得巨大成功,其性能依赖于海量文本数据的训练,包括受版权保护内容和用户生成知识。这带来了隐私泄露和版权侵犯风险。为保障用户“被遗忘权”,机器遗忘成为研究热点——在不损害模型预测性能的前提下,移除特定训练样本的信息。由于语言模型的黑箱特性,现有方法多关注输出层面的影响缓解,未考虑隐空间中样本的几何分布。为此,我们提出深度对比遗忘框架DeepCUT,直接优化模型隐空间以实现遗忘。在真实数据集上的实验表明,DeepCUT在有效性与效率上均显著优于基线方法。

原文摘要 · Abstract (English)

The past a few years have witnessed the great success of large language models, demonstrating powerful capabilities in comprehending textual data and generating human-like languages. Large language models achieve success by being trained on vast amounts of textual data, including online sources with copyrighted content and user-generated knowledge. However, this comes at a cost: the potential risk of exposing users' privacy and violating copyright protections. Thus, to safeguard individuals' "right to be forgotten", there has been increasing interests in machine unlearning -- the process of removing information carried by particular training samples from a model while not deteriorating its predictive quality. This is a challenging task due to the black-box nature of language models. Most existing studies focus on mitigating the impact of those forgot samples upon a model's outputs, and do not explicitly consider the geometric distributions of samples in the latent space of a model. To address this issue, we propose a machine unlearning framework, named Deep Contrastive Unlearning for fine-Tuning (DeepCUT) language models. Our proposed model achieves machine unlearning by directly optimizing the latent space of a model. Comprehensive experiments on real-world datasets demonstrate the effectiveness and efficiency of DeepCUT with consistent and significant improvement over baseline methods.

大模型机器遗忘隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。