arXiv:2510.25117cs.CL2025-10综述被引 4

系统梳理180篇大模型遗忘技术论文,解决敏感信息泄露问题。

A Survey on Unlearning in Large Language Models

  • 按干预阶段与参数修改/选择分类,构建新评估框架。
  • 对比18个数据集、10类记忆度指标,全面评估效果。
  • 适合关注隐私安全与合规的大模型研究者参考。

大型语言模型(LLMs)虽具备强大能力,但其在海量语料上训练可能带来敏感信息记忆风险。为缓解此问题并符合法律要求,遗忘技术成为关键手段,可在不损害整体性能前提下选择性擦除特定知识。本综述系统回顾了2021年以来超过180篇关于大模型遗忘的论文。首先提出一种基于干预阶段的新分类法,区分参数修改与参数选择策略,深化理解并支持更精准比较。其次,从多维度分析评估范式:针对18个现有基准,从任务格式、内容和实验设计角度进行对比,提供实践指导;对记忆度指标,将其细分为10类,分析各自优劣与适用场景,并综述模型效用、鲁棒性与效率相关评估方法。通过探讨当前挑战与未来方向,旨在推动大模型遗忘技术发展及安全AI体系建设。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate remarkable capabilities, but their training on massive corpora poses significant risks from memorized sensitive information. To mitigate these issues and align with legal standards, unlearning has emerged as a critical technique to selectively erase specific knowledge from LLMs without compromising their overall performance. This survey provides a systematic review of over 180 papers on LLM unlearning published since 2021. First, it introduces a novel taxonomy that categorizes unlearning methods based on the phase in the LLM pipeline of the intervention. This framework further distinguishes between parameter modification and parameter selection strategies, thus enabling deeper insights and more informed comparative analysis. Second, it offers a multidimensional analysis of evaluation paradigms. For datasets, we compare 18 existing benchmarks from the perspectives of task format, content, and experimental paradigms to offer actionable guidance. For metrics, we move beyond mere enumeration by dividing knowledge memorization metrics into 10 categories to analyze their advantages and applicability, while also reviewing metrics for model utility, robustness, and efficiency. By discussing current challenges and future directions, this survey aims to advance the field of LLM unlearning and the development of secure AI systems.

大模型遗忘技术隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。