让大模型删除特定知识,避免隐私与版权风险。
Unlearning in LLMs: Methods, Evaluation, and Open Challenges
- 按数据、参数、架构等策略分类,设计可删知识的模型方法
- 构建评测基准,量化遗忘效果与知识保留程度
- 适合关注模型安全与合规的研究者和开发者
大型语言模型在自然语言处理任务中取得显著成功,但其广泛应用引发隐私、版权、安全和偏见等关切。机器遗忘作为一种新兴范式,可在不重新训练的情况下选择性移除模型中的知识或数据。本文系统综述了大模型的遗忘方法,将其分为数据导向、参数导向、架构导向、混合及其他策略。同时梳理了评估体系,包括用于衡量遗忘效果、知识保留和鲁棒性的基准、指标与数据集。最后指出关键挑战:可扩展效率、形式化保证、跨语言与多模态遗忘,以及对抗性重学的鲁棒性。本文旨在为发展可靠且负责任的大模型遗忘技术提供路线图。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable success across natural language processing tasks, yet their widespread deployment raises pressing concerns around privacy, copyright, security, and bias. Machine unlearning has emerged as a promising paradigm for selectively removing knowledge or data from trained models without full retraining. In this survey, we provide a structured overview of unlearning methods for LLMs, categorizing existing approaches into data-centric, parameter-centric, architecture-centric, hybrid, and other strategies. We also review the evaluation ecosystem, including benchmarks, metrics, and datasets designed to measure forgetting effectiveness, knowledge retention, and robustness. Finally, we outline key challenges and open problems, such as scalable efficiency, formal guarantees, cross-language and multimodal unlearning, and robustness against adversarial relearning. By synthesizing current progress and highlighting open directions, this paper aims to serve as a roadmap for developing reliable and responsible unlearning techniques in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。