用不可察觉的干扰保护图像隐私,防止VLM提取敏感信息
Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection

- 结合扩散模型与可逆网络生成无损恢复的隐蔽对抗样本
- 在多个数据集和VLM上实现高保真度与强迁移性
- 适合关注多模态隐私安全的研究者与应用开发者
视觉语言模型(VLMs)虽具备强大多模态能力,却易受文本驱动的隐私攻击:攻击者通过爬取网络图片并查询VLM以提取敏感属性。现有可逆对抗样本方法仅适用于纯视觉任务,在多模态场景中失效;而当前针对VLM的对抗样本依赖高频噪声,严重损害视觉质量。本文提出CloakDiff,首个面向VLM的可逆、高保真隐私保护框架。该方法通过扩散式对抗编辑结合可逆网络,嵌入原始图像以实现无损恢复;同时扰动像素空间嵌入并操控潜在交叉注意力图,确保跨模型与跨提示的强迁移性,同时保持全局视觉结构。为进一步提升保真度,设计了基于EDM的启发式采样策略,用于对抗引导的扩散过程。在多个数据集和VLM上的实验表明,CloakDiff实现了多模态隐私保护、高视觉质量和可逆性。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) offer powerful multimodal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract sensitive attributes. Existing reversible adversarial example (RAE) methods protect images in purely visual tasks but fail in multimodal settings, and current adversarial examples on VLMs rely on high frequency noise that severely degrades visual quality. We propose CloakDiff, the first framework for reversible, high fidelity privacy protection against text-based query attacks in VLMs. CloakDiff produces imperceptible adversarial examples by combining diffusion based adversarial editing with an invertible network that embeds the original image for lossless recovery. It perturbs both pixel space embeddings and manipulates latent cross attention maps to ensure strong cross-model and cross-prompt transferability while preserving global visual structure. To further enhance fidelity, we design EDM Heuristic Sampling, a principled diffusion schedule for adversarial guidance. Experiments on multiple datasets and VLMs demonstrate that CloakDiff delivers multimodal privacy preservation with high visual quality and reversibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。