arXiv:2506.12618cs.CL2025-06NeurIPS被引 74

构建统一框架加速大模型数据遗忘研究

OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics

  • 设计标准化框架OpenUnlearning,集成13种遗忘方法与16项评估
  • 在3个基准上测试超450个检查点,发现多数方法遗忘不彻底
  • 提出新评测标准,可检验评估指标本身是否可靠,适合安全研究者

可靠的模型遗忘对保障大语言模型在数据隐私、安全和合规场景下的部署至关重要。然而,该任务面临核心挑战:难以可靠衡量遗忘是否真正发生。当前方法分散、评估指标不一,导致比较困难且难以复现。为此,我们提出OpenUnlearning,一个专为大模型遗忘方法与评估指标统一评测而设计的标准化、可扩展框架。该框架整合了13种遗忘算法和16种多样评估,覆盖3个主流基准(TOFU、MUSE、WMDP),并公开发布450多个检查点以支持遗忘行为分析。基于此,我们提出一个新型元评估基准,专门用于检验评估指标的忠实性与鲁棒性。同时,我们对多种遗忘方法进行了全面测评,并提供系统性对比分析。整体上,本工作为大模型遗忘研究建立了清晰、社区驱动的严谨发展路径。

原文摘要 · Abstract (English)

Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured. Yet the task is inherently challenging, partly due to difficulties in reliably measuring whether unlearning has truly occurred. Moreover, fragmentation in current methodologies and inconsistent evaluation metrics hinder comparative analysis and reproducibility. To unify and accelerate research efforts, we introduce OpenUnlearning, a standardized and extensible framework designed explicitly for benchmarking both LLM unlearning methods and metrics. OpenUnlearning integrates 13 unlearning algorithms and 16 diverse evaluations across 3 leading benchmarks (TOFU, MUSE, and WMDP) and also enables analyses of forgetting behaviors across 450+ checkpoints we publicly release. Leveraging OpenUnlearning, we propose a novel meta-evaluation benchmark focused specifically on assessing the faithfulness and robustness of evaluation metrics themselves. We also benchmark diverse unlearning methods and provide a comparative analysis against an extensive evaluation suite. Overall, we establish a clear, community-driven pathway toward rigorous development in LLM unlearning research.

大模型遗忘学习评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。