arXiv:2506.02541cs.LGcs.AI2025-06ACL

提出新方法让大模型删掉隐私信息后还能给出有信息量的合理回答。

Rethinking Post-Unlearning Behavior of Large Vision-Language Models

  • 设计新任务,要求删除隐私后仍能提供视觉相关且有用的回答。
  • 实验显示现有方法删完会瞎编或拒绝回答,新方法避免这些问题。
  • 适合关注大模型隐私安全与生成质量平衡的研究者使用。

大型视觉语言模型(LVLMs)能够识别图像中的人物并泄露其敏感个人信息,引发严重隐私问题。机器遗忘旨在从模型中移除此类知识,但现有方法很少规定遗忘后模型应输出什么,导致出现遗忘后遗症:退化、幻觉或过度拒绝回应。我们指出,对于生成式LVLMs而言,关注遗忘后输出的质量与信息量至关重要,不应仅依赖简单的抑制策略。为此,我们提出一种新的LVLM遗忘任务,要求模型在保护隐私的同时提供具有信息量且视觉一致的回应。我们还提出了PUBG这一新遗忘方法,显式引导遗忘后的输出分布趋向理想状态。实验表明,尽管现有方法虽能防止隐私泄露,但普遍存在遗忘后遗症;而PUBG有效缓解了这些问题,在不泄露遗忘目标隐私的前提下,生成了视觉一致且富有信息量的回应。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) can recognize individuals in images and disclose sensitive personal information about them, raising critical privacy concerns. Machine unlearning aims to remove such knowledge from the model. However, existing methods rarely prescribe what the model should output in place of the forgotten content, leading to Unlearning Aftermaths: degenerate, hallucinated, or excessively refused responses. We argue that, especially for generative LVLMs, it is crucial to consider the quality and informativeness of post-unlearning responses rather than relying solely on naive suppression. To address this, we introduce a new unlearning task for LVLMs that requires models to provide privacy-preserving yet informative and visually grounded responses. We also propose PUBG, a novel unlearning method that explicitly guides post-unlearning behavior toward a desirable output distribution. Experiments show that, while existing methods suffer from Unlearning Aftermaths despite successfully preventing privacy violations, PUBG effectively mitigates these issues, generating visually grounded and informative responses without privacy leakage for forgotten targets.

视觉语言模型隐私保护机器遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。