为多模态大模型设计新评估基准,解决隐私删除中的概念干扰问题。
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
- 构建虚构人物与事件场景数据集,评估多模态模型删减能力。
- 发现删除一个概念会损害相关概念性能,存在跨概念干扰现象。
- 提出缓解冲突目标的新方法,适合关注模型隐私安全的研究者。
多模态大语言模型(MLLM)在视觉-语言任务中表现卓越,但其依赖海量互联网数据引发严重的隐私与安全问题。机器遗忘(MU)作为关键解决方案,可在不重新训练的情况下选择性移除特定信息。然而,当前对MLLM的机器遗忘评估仍不充分,现有基准往往只关注实体,忽视更广泛的视觉概念及概念间的语义耦合。为此,我们提出PEBench——一个新型基准,包含虚构个人实体及其对应事件场景,用于全面评估MLLM中的机器遗忘效果。利用该基准评估五种MU方法,揭示其优劣。研究发现,删除某一概念可能意外降低同一图像中相关概念的性能,这一现象称为跨概念干扰。此外,同时删除人物与事件概念存在困难,并提出有效方法缓解此类冲突目标。源代码与基准已公开于https://pebench.github.io。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved remarkable success in vision-language tasks, but their reliance on vast, internet-sourced data raises significant privacy and security concerns. Machine unlearning (MU) has emerged as a critical technique to address these issues, enabling the selective removal of targeted information from pre-trained models without costly retraining. However, the evaluation of MU for MLLMs remains inadequate. Existing benchmarks often lack a comprehensive scope, focusing narrowly on entities while overlooking the unlearning of broader visual concepts and the inherent semantic coupling between them. To bridge this gap, we introduce, PEBench, a novel benchmark designed to facilitate a thorough assessment of MU in MLLMs. PEBench features a fictitious dataset of personal entities and corresponding event scenes to evaluate unlearning across these distinct yet entangled concepts. We leverage this benchmark to evaluate five MU methods, revealing their unique strengths and weaknesses. Our findings show that unlearning one concept can unintentionally degrade performance on related concepts within the same image, a challenge we term cross-concept interference. Furthermore, we demonstrate the difficulty of unlearning person and event concepts simultaneously and propose an effective method to mitigate these conflicting objectives. The source code and benchmark are publicly available at https://pebench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。