arXiv:2510.22535cs.AIcs.CL2025-10ACL被引 7

构建足球转会谣言数据集,评估多模态大模型删去错误信息的能力

OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models

  • 基于足球转会谣言设计多模态数据集,支持选择性删去文本或图像知识
  • 发现仅删除文本无法消除多模态谣言,且删后易被恢复,存在提示攻击漏洞
  • 适用于研究隐私保护与虚假信息清除的AI研究人员

多模态大语言模型(MLLMs)的发展加剧了数据隐私担忧,机器删忆(MU)——选择性删除已学信息——成为迫切需求。然而现有MLLM删忆基准受限于图像多样性不足、潜在误差及评估场景匮乏,难以反映真实应用复杂性。为此,我们提出OFFSIDE,一个基于足球转会谣言的多模态虚假信息删忆基准。该手动构建的数据集包含80名球员的15.68万条记录,提供四个测试集,用于评估遗忘效果、泛化能力、实用性与鲁棒性。支持选择性删忆与修正重学等高级设置,尤其涵盖单模态删忆(仅删除文本知识)。对多个基线的广泛评估揭示:(1) 仅删文本的方法在多模态谣言中失效;(2) 忘记效果主要由灾难性遗忘驱动;(3) 所有方法均难以应对‘视觉谣言’(谣言出现在图像中);(4) 删去的谣言可轻易恢复;(5) 所有方法均易受提示攻击。结果暴露当前方法重大缺陷,亟需更鲁棒的多模态删忆方案。代码已开源:https://github.com/zh121800/OFFSIDE

原文摘要 · Abstract (English)

Advances in Multimodal Large Language Models (MLLMs) intensify concerns about data privacy, making Machine Unlearning (MU), the selective removal of learned information, a critical necessity. However, existing MU benchmarks for MLLMs are limited by a lack of image diversity, potential inaccuracies, and insufficient evaluation scenarios, which fail to capture the complexity of real-world applications. To facilitate the development of MLLMs unlearning and alleviate the aforementioned limitations, we introduce OFFSIDE, a novel benchmark for evaluating misinformation unlearning in MLLMs based on football transfer rumors. This manually curated dataset contains 15.68K records for 80 players, providing a comprehensive framework with four test sets to assess forgetting efficacy, generalization, utility, and robustness. OFFSIDE supports advanced settings like selective unlearning and corrective relearning, and crucially, unimodal unlearning (forgetting only text data). Our extensive evaluation of multiple baselines reveals key findings: (1) Unimodal methods (erasing text-based knowledge) fail on multimodal rumors; (2) Unlearning efficacy is largely driven by catastrophic forgetting; (3) All methods struggle with "visual rumors" (rumors appear in the image); (4) The unlearned rumors can be easily recovered and (5) All methods are vulnerable to prompt attacks. These results expose significant vulnerabilities in current approaches, highlighting the need for more robust multimodal unlearning solutions. The code is available at https://github.com/zh121800/OFFSIDE

多模态删忆虚假信息隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。