arXiv:2410.22108cs.CLcs.AI2024-10NAACL被引 71

首个针对多模态大模型的隐私遗忘基准,评估数据删除效果。

Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench

  • 构建包含500个虚构与153个名人资料的多模态遗忘测试集
  • 发现单模态遗忘在生成任务中表现更好,多模态遗忘在分类任务中更优
  • 为多模态大模型隐私保护提供可量化评估工具,适合研究者参考

生成式模型如大语言模型(LLM)和多模态大语言模型(MLLM)在大规模网络语料上训练时可能记忆并泄露个人敏感数据,引发法律与伦理问题。尽管已有工作针对LLM的机器遗忘进行研究,但对MLLM的探索仍较少。为此,我们提出多模态大语言模型遗忘基准(MLLMU-Bench),旨在推动多模态机器遗忘的理解。该基准包含500个虚构人物与153个公众名人资料,每份资料含超过14个定制问答对,从多模态(图像+文本)与单模态(仅文本)角度进行评估。基准分为四个数据集,用于衡量遗忘算法的有效性、泛化能力与模型实用性。我们还提供了基于现有生成模型遗忘算法的基线结果。令人意外的是,实验显示单模态遗忘算法在生成与完形填空任务中表现更佳,而多模态遗忘方法在处理多模态输入的分类任务中更具优势。

原文摘要 · Abstract (English)

Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we introduce Multimodal Large Language Model Unlearning Benchmark (MLLMU-Bench), a novel benchmark aimed at advancing the understanding of multimodal machine unlearning. MLLMU-Bench consists of 500 fictitious profiles and 153 profiles for public celebrities, each profile feature over 14 customized question-answer pairs, evaluated from both multimodal (image+text) and unimodal (text) perspectives. The benchmark is divided into four sets to assess unlearning algorithms in terms of efficacy, generalizability, and model utility. Finally, we provide baseline results using existing generative model unlearning algorithms. Surprisingly, our experiments show that unimodal unlearning algorithms excel in generation and cloze tasks, while multimodal unlearning approaches perform better in classification tasks with multimodal inputs.

多模态模型隐私保护机器遗忘基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。