arXiv:2504.12681cs.CLcs.AI2025-04中稿 · IJCNN 2025被引 4

提出GRAIL框架,精准删去大模型中隐私与版权信息而不伤核心能力

GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs

  • 基于多领域梯度信息定位需删除与需保留的参数
  • 在多领域场景下实现90%以上知识删除且性能仅降17%
  • 适合需要合规删改的AI公司和监管机构使用

大规模语言模型在训练中常学习到敏感信息,引发隐私与版权法律风险。重新训练整个模型成本过高,且现有单领域遗忘方法无法处理跨领域知识交织问题,导致过度删减或性能下降。为此,我们提出GRAIL(GRadient-based AdaptIve unLearning)框架,利用多领域梯度信息精确区分需遗忘与需保留的参数,并采用参数级自适应定位策略,选择性移除特定知识同时保护各领域关键参数。在多个遗忘基准测试中,GRAIL在遗忘效果上达到现有方法水平,同时相比最优前序方法,知识保留成功率提升最高达17%。研究为大规模预训练语言模型中敏感信息的有效管理提供了新范式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) trained on extensive datasets often learn sensitive information, which raises significant social and legal concerns under principles such as the "Right to be forgotten." Retraining entire models from scratch to remove undesired information is both costly and impractical. Furthermore, existing single-domain unlearning methods fail to address multi-domain scenarios, where knowledge is interwoven across domains such as privacy and copyright, creating overlapping representations that lead to excessive knowledge removal or degraded performance. To tackle these issues, we propose GRAIL (GRadient-based AdaptIve unLearning), a novel multi-domain unlearning framework. GRAIL leverages gradient information from multiple domains to precisely distinguish the unlearning scope from the retention scope, and applies an adaptive parameter-wise localization strategy to selectively remove targeted knowledge while preserving critical parameters for each domain. Experimental results on unlearning benchmarks show that GRAIL achieves unlearning success on par with the existing approaches, while also demonstrating up to 17% stronger knowledge retention success compared to the previous state-of-art method. Our findings establish a new paradigm for effectively managing and regulating sensitive information in large-scale pre-trained language models.

大模型遗忘隐私保护知识保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。