arXiv:2503.01854cs.CLcs.AI2025-03综述被引 37

系统梳理大模型数据删除技术,解决敏感信息遗忘难题

A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models

  • 按方法原理分类现有大模型遗忘技术,构建完整框架
  • 总结评估指标与基准测试,统一研究评价标准
  • 适合关注隐私保护与合规的大模型研发者参考

本研究探讨大语言模型(LLMs)中的机器遗忘技术,即 extit{LLM unlearning}。该技术为消除模型中不良数据(如敏感或违法信息)的影响提供了一种规范方法,可在不重新训练的前提下保留模型整体性能。尽管研究兴趣日益增长,但尚无系统性综述对现有工作进行归纳并提炼关键洞见。本文旨在填补这一空白:首先定义并阐述LLM遗忘的范式,随后构建全面的分类体系;接着对现有遗忘方法进行归类,分析其优缺点;此外,回顾评估指标与基准测试,提供结构化的方法论概述;最后,提出未来研究的潜在方向,指出领域内的关键挑战与机遇。

原文摘要 · Abstract (English)

This study investigates the machine unlearning techniques within the context of large language models (LLMs), referred to as \textit{LLM unlearning}. LLM unlearning offers a principled approach to removing the influence of undesirable data (e.g., sensitive or illegal information) from LLMs, while preserving their overall utility without requiring full retraining. Despite growing research interest, there is no comprehensive survey that systematically organizes existing work and distills key insights; here, we aim to bridge this gap. We begin by introducing the definition and the paradigms of LLM unlearning, followed by a comprehensive taxonomy of existing unlearning studies. Next, we categorize current unlearning approaches, summarizing their strengths and limitations. Additionally, we review evaluation metrics and benchmarks, providing a structured overview of current assessment methodologies. Finally, we outline promising directions for future research, highlighting key challenges and opportunities in the field.

大模型隐私保护机器遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。