arXiv:2506.17265cs.LGcs.AI2025-06EMNLP被引 5

提出隐蔽攻击方法,让遗忘敏感信息的多模态模型重新泄露数据。

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

  • 设计通用噪声模式,通过嵌入空间对齐提升隐蔽性。
  • 单次训练的噪声可复现未学习内容,且在未见图像上有效。
  • 揭示模型遗忘非真正删除,适合隐私安全研究者参考。

多模态大语言模型(MLLMs)在海量数据上训练可能记忆敏感个人信息和图像,带来严重隐私风险。为缓解此问题,现有遗忘方法通过微调使模型减少对敏感信息的响应。然而,尚不清楚知识是真正遗忘还是仅被隐藏。为此,本文提出新型大语言模型遗忘攻击问题,旨在恢复已被遗忘模型中的知识。我们提出隐蔽遗忘攻击(SUA)框架,学习一种通用噪声模式;当应用于输入图像时,该噪声可触发模型重现被遗忘内容。尽管像素级扰动视觉上细微,但在语义嵌入空间中仍可被检测,易遭防御。为此,我们引入嵌入对齐损失,最小化扰动前后图像嵌入差异,确保攻击在语义上不可察觉。实验表明,SUA能有效从已遗忘的MLLM中恢复敏感信息。此外,所学噪声具有强泛化能力:仅用部分样本训练的单个扰动即可在未见图像中揭示遗忘内容,表明知识重现并非偶然失败,而是持续行为。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing serious privacy risks. To mitigate this, MLLM unlearning methods are proposed, which fine-tune MLLMs to reduce the ``forget'' sensitive information. However, it remains unclear whether the knowledge has been truly forgotten or just hidden in the model. Therefore, we propose to study a novel problem of LLM unlearning attack, which aims to recover the unlearned knowledge of an unlearned LLM. To achieve the goal, we propose a novel framework Stealthy Unlearning Attack (SUA) framework that learns a universal noise pattern. When applied to input images, this noise can trigger the model to reveal unlearned content. While pixel-level perturbations may be visually subtle, they can be detected in the semantic embedding space, making such attacks vulnerable to potential defenses. To improve stealthiness, we introduce an embedding alignment loss that minimizes the difference between the perturbed and denoised image embeddings, ensuring the attack is semantically unnoticeable. Experimental results show that SUA can effectively recover unlearned information from MLLMs. Furthermore, the learned noise generalizes well: a single perturbation trained on a subset of samples can reveal forgotten content in unseen images. This indicates that knowledge reappearance is not an occasional failure, but a consistent behavior.

隐私安全模型攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。