arXiv:2607.06649cs.CRcs.LG2026-07

提出新攻击方法,从声称已遗忘隐私信息的多模态模型中恢复敏感内容。

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

  • 通过优化提示后缀诱导模型生成私有内容
  • 在多个基准上实现近完全恢复被删除的敏感信息
  • 揭示现有遗忘机制的根本脆弱性,适合安全研究者参考

多模态大语言模型(MLLMs)通过联合训练大规模文本与视觉数据,在跨模态任务上表现优异,但可能无意中编码隐私敏感样本,引发隐私或版权争议。为此,多模态机器遗忘(MMU)被提出以有效使模型遗忘私有信息。然而,当模型公开后,现有遗忘方法的鲁棒性未被充分验证。本文提出一种新型对抗性策略——提示优化参数抖动(POPS),旨在从已执行遗忘的MLLM中恢复本应被删除的多模态知识。该方法通过提示后缀优化诱导目标MLLM生成潜在私有样本,并利用这些合成输出进行微调,从而暴露真实私有信息。在多个MMU基准上的实验表明,现有遗忘算法存在显著弱点。POPS甚至可在遗忘后的模型上实现近乎完全的信息恢复,暴露出基于代表性MMU的隐私保护机制的根本性漏洞。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and visual data, where privacy-sensitive examples could be unintentionally encoded, raising concerns about privacy or copyright violation. To this end, Multi-modality Machine Unlearning (MMU) was proposed as a mitigation that can effectively force MLLMs to forget private information. However, the robustness of such unlearning methods is not fully exploited when the model is published and accessible to malicious users. In this paper, we propose a novel adversarial strategy, namely Prompt-Optimized Parameter Shaking (POPS), aiming to recover the supposedly unlearned multi-modality knowledge from the MLLMs. Our method elicits the victim MLLMs to generate potential private examples via prompt-suffix optimization, and then exploits these synthesized outputs to fine-tune the models so they disclose the true private information. The experiments on the different MMU benchmarks reveal substantial weaknesses in the existing MMU algorithms. Our POPS can even achieve a near-complete recovery of supposedly erased sensitive information on the unlearned MLLMs, exposing fundamental vulnerabilities that challenge the foundational robustness of representative MMU-based privacy protections.

多模态机器遗忘隐私安全对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。