为多模态大模型设计真实遗忘评估框架,发现现有方法难删预训练知识且不支持分步遗忘。
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
- 提出双维度评估框架:按知识获取阶段区分遗忘效果,检验长期分步遗忘能力。
- 实测表明:多数方法能删微调知识,但无法清除预训练阶段学习的信息。
- 单次批量遗忘有效的模型,分步多次遗忘时性能显著下降,暴露实际应用缺陷。
近年来,遗忘技术作为应对大语言模型(LLMs)和大多模态模型(LMMs)隐私与版权问题的手段受到关注。尽管已有针对LLMs的遗忘基准,但面向LMMs的实用评估框架仍不充分。现有LMM遗忘基准仅考虑通过单一操作消除微调知识的场景。本文提出PULSE协议,引入两个关键视角:(i) 预训练知识遗忘,分析不同知识获取阶段的遗忘效果;(ii) 长期可持续性评估,应对连续遗忘请求。我们在此框架下评估现有遗忘方法。结果表明:虽部分技术可有效删除微调知识,却难以消除预训练阶段习得的信息;且在单次批量遗忘中表现良好的方法,在分步处理相同数据时性能大幅下降。
原文摘要 · Abstract (English)
In recent years, unlearning techniques, which are methods for inducing a model to "forget" previously learned information, have attracted attention as a way to address privacy and copyright concerns in large language models (LLMs) and large multimodal models (LMMs). While several unlearning benchmarks have been established for LLMs, a practical evaluation framework for unlearning in LMMs has been less explored. Specifically, existing unlearning benchmark for LMMs considers only scenarios in which the model is required to unlearn fine-tuned knowledge through a single unlearning operation. In this study, we introduce PULSE protocol for realistic unlearning scenarios for LMMs by introducing two critical perspectives: (i) Pre-trained knowledge Unlearning for analyzing the effect across different knowledge acquisition phases and (ii) Long-term Sustainability Evaluation to address sequential requests. We then evaluate existing unlearning methods along these dimensions. Our results reveal that, although some techniques can successfully unlearn knowledge acquired through fine-tuning, they struggle to eliminate information learned during pre-training. Moreover, methods that effectively unlearn a batch of target data in a single operation exhibit substantial performance degradation when the same data are split and unlearned sequentially.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。