提出无需微调的隐私保护方法,有效阻止多模态模型泄露个人身份信息。
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
- 通过概念引导识别并修改模型内部与敏感信息相关的状态。
- 在各类隐私任务上平均拒绝率达93.3%,对其他任务影响小。
- 构建真实场景模拟数据集,推动多模态隐私研究发展。
多模态大语言模型在处理多种模态数据方面展现出强大能力,但其先进性能也引发严重的隐私风险,尤其涉及个人身份信息(PII)泄露问题。尽管已有针对单模态语言模型的研究,但多模态场景下的漏洞尚未被充分探索。本文聚焦视觉语言模型(VLMs),这类模型涵盖最易泄露PII的视觉与文本两种模态。我们提出一种概念引导的缓解方法,识别并修改模型内部与PII相关内容相关的状态,使VLM能有效且高效地拒绝敏感任务,无需重新训练或微调。同时,针对当前缺乏多模态PII数据集的问题,我们构建了多个模拟真实场景的数据集。实验表明,该方法在多种PII相关任务上平均拒绝率达到93.3%,对无关任务性能影响极小。我们进一步在不同条件下评估其表现,验证了方法的适应性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing and reasoning over diverse modalities, but their advanced abilities also raise significant privacy concerns, particularly regarding Personally Identifiable Information (PII) leakage. While relevant research has been conducted on single-modal language models to some extent, the vulnerabilities in the multimodal setting have yet to be fully investigated. In this work, we investigate these emerging risks with a focus on vision language models (VLMs), a representative subclass of MLLMs that covers the two modalities most relevant for PII leakage, vision and text. We introduce a concept-guided mitigation approach that identifies and modifies the model's internal states associated with PII-related content. Our method guides VLMs to refuse PII-sensitive tasks effectively and efficiently, without requiring re-training or fine-tuning. We also address the current lack of multimodal PII datasets by constructing various ones that simulate real-world scenarios. Experimental results demonstrate that the method can achieve an average refusal rate of 93.3% for various PII-related tasks with minimal impact on unrelated model performances. We further examine the mitigation's performance under various conditions to show the adaptability of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。