arXiv:2507.01271cs.LGcs.AI2025-07中稿 · NeurIPS被引 5

为多模态大模型设计真实遗忘评估框架,发现现有方法难删预训练知识且不支持分步遗忘。

PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning

  • 提出双维度评估框架:按知识获取阶段区分遗忘效果,检验长期分步遗忘能力。
  • 实测表明:多数方法能删微调知识,但无法清除预训练阶段学习的信息。
  • 单次批量遗忘有效的模型,分步多次遗忘时性能显著下降,暴露实际应用缺陷。

近年来,遗忘技术作为应对大语言模型(LLMs)和大多模态模型(LMMs)隐私与版权问题的手段受到关注。尽管已有针对LLMs的遗忘基准,但面向LMMs的实用评估框架仍不充分。现有LMM遗忘基准仅考虑通过单一操作消除微调知识的场景。本文提出PULSE协议,引入两个关键视角:(i) 预训练知识遗忘,分析不同知识获取阶段的遗忘效果;(ii) 长期可持续性评估,应对连续遗忘请求。我们在此框架下评估现有遗忘方法。结果表明:虽部分技术可有效删除微调知识,却难以消除预训练阶段习得的信息;且在单次批量遗忘中表现良好的方法,在分步处理相同数据时性能大幅下降。

原文摘要 · Abstract (English)

In recent years, unlearning techniques, which are methods for inducing a model to "forget" previously learned information, have attracted attention as a way to address privacy and copyright concerns in large language models (LLMs) and large multimodal models (LMMs). While several unlearning benchmarks have been established for LLMs, a practical evaluation framework for unlearning in LMMs has been less explored. Specifically, existing unlearning benchmark for LMMs considers only scenarios in which the model is required to unlearn fine-tuned knowledge through a single unlearning operation. In this study, we introduce PULSE protocol for realistic unlearning scenarios for LMMs by introducing two critical perspectives: (i) Pre-trained knowledge Unlearning for analyzing the effect across different knowledge acquisition phases and (ii) Long-term Sustainability Evaluation to address sequential requests. We then evaluate existing unlearning methods along these dimensions. Our results reveal that, although some techniques can successfully unlearn knowledge acquired through fine-tuning, they struggle to eliminate information learned during pre-training. Moreover, methods that effectively unlearn a batch of target data in a single operation exhibit substantial performance degradation when the same data are split and unlearned sequentially.

模型遗忘多模态模型隐私保护评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。