保护视觉隐私,让大模型在不泄露图像信息下正常工作
When Visual Privacy Protection Meets Multimodal Large Language Models
- 设计基于帕累托最优的优化目标,平衡隐私与模型性能
- 无需模型内部信息,仅通过输入输出实现隐私保护
- 适用于无法访问模型结构的黑箱大模型服务场景
多模态大语言模型(MLLM)及其云端服务(如GPT-4V)的普及引发了视觉数据隐私泄露的担忧。由于模型部署于云端,用户需上传图像和视频,存在严重隐私风险。然而,如何应对这一问题仍缺乏研究。本文针对黑箱场景(仅可访问输入输出,不可知内部结构)提出新框架:通过帕累托最优设计学习目标,实现隐私保护与模型性能的更好权衡,并引入关键历史增强优化方法,有效训练该框架。实验表明,该方法在多个基准上均具有效性。
原文摘要 · Abstract (English)
The emergence of Multimodal Large Language Models (MLLMs) and the widespread usage of MLLM cloud services such as GPT-4V raised great concerns about privacy leakage in visual data. As these models are typically deployed in cloud services, users are required to submit their images and videos, posing serious privacy risks. However, how to tackle such privacy concerns is an under-explored problem. Thus, in this paper, we aim to conduct a new investigation to protect visual privacy when enjoying the convenience brought by MLLM services. We address the practical case where the MLLM is a "black box", i.e., we only have access to its input and output without knowing its internal model information. To tackle such a challenging yet demanding problem, we propose a novel framework, in which we carefully design the learning objective with Pareto optimality to seek a better trade-off between visual privacy and MLLM's performance, and propose critical-history enhanced optimization to effectively optimize the framework with the black-box MLLM. Our experiments show that our method is effective on different benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。