提出无需训练的隐私保护方法,让多模态模型看清世界又不泄露敏感信息。
Seeing Without Exposing: Adaptive Privacy Control for Open-World, Context-Hungry MLLMs

- 通过语义等价替换实现隐私元素漂移,同时锚定上下文信息保持图像可用性。
- 在4个主流多模态模型上平均提升10.4%文本类隐私保护效果,8.5%上下文保留率。
- 构建覆盖22类隐私的AdaptShield评测基准,兼顾隐私与内容实用性评估。
多模态大语言模型(MLLM)带来了新的隐私挑战:用户输入中可能包含不可预测的敏感信息,而模型推理又依赖丰富的视觉上下文,这些上下文本身也可能涉及隐私。现有方法依赖预定义敏感类别和固定模糊策略,难以应对MLLM场景下的复杂需求。为此,我们提出无需训练的锚定隐私漂移(APD)方法,将敏感元素漂移到语义等价的替代项,同时锚定上下文线索以保留源图像信息。为系统评估隐私保护与上下文保留的双重目标,我们构建了AdaptShield基准,涵盖22类隐私,结合传统隐私指标与基于MLLM的上下文效用评估。大量实验表明,该方法在四个MLLM系列(Qwen2.5、Qwen3、InternVL3、InternVL3.5)上实现了平衡提升,文本类隐私保护平均增益10.4%,基于MLLM的评估平均提升8.5%。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have raised new privacy challenges. On the data side, user-provided inputs often include unpredictable sensitive information; while on the downstream task side, model reasoning depends on rich visual context that may itself be privacy-sensitive. Existing privacy protection methods, however, rely on predefined sensitive categories and fixed obfuscation strategies, struggling to tackle such challenges in MLLMs. To address this dilemma, we propose Anchored Privacy Drifting (APD), a training-free method that drifts privacy-sensitive elements toward semantically equivalent alternatives while anchoring contextual cues to the source image. To systematically evaluate this dual objective of privacy protection and contextual preservation, we introduce AdaptShield, a comprehensive benchmark covering 22 privacy categories, which combines conventional privacy metrics with MLLM-based assessments of contextual utility. Extensive experiments show that our method achieves balanced improvements in both privacy sanitization and content retention, with average gains of 10.4% on textual categories and 8.5% under MLLM-based evaluation across four MLLM series, i.e., Qwen2.5, Qwen3, InternVL3, and InternVL3.5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。