首个隐私保护多模态数据集,专治大模型泄露人脸生物特征
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
- 从原始数据中系统清除显式和隐式生物特征信息,构建隐私安全数据集
- 发现主流模型在未被提问时仍会泄露种族、性别等敏感信息
- 提供新评测基准,适合关注模型隐私安全的研究者与开发者
多模态大语言模型在视觉-语言任务中表现卓越,但常会在未被明确要求的情况下推断并暴露种族、性别、年龄、体脂率、眼色等敏感生物特征,引发严重隐私担忧。尽管关注度提升,目前尚无公开数据集或评测基准用于全面评估或缓解此类生物特征泄露问题。为此,我们提出PRISM(隐私感知敏感模态响应评估)基准,用于评估模型在两类场景下的表现:(1)拒绝生物特征相关查询;(2)在一般回答中避免隐性生物特征泄露,同时保持语义忠实度。我们对广泛使用的LLaVA数据集进行审计,发现预训练与指令数据中存在广泛的生物特征泄露。为此,我们构建了首个隐私保护型多模态训练数据集Safe-LLaVA,通过系统移除显式与隐式生物特征信息而成。在PRISM上的评估显示,不同模型对各类生物特征均存在显著泄露,揭示了详细的隐私风险。我们基于Safe-LLaVA微调模型,证明其能大幅降低生物特征泄露。Safe-LLaVA与PRISM共同确立了多模态大模型隐私对齐开发与评估的新标准。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks. However, these models often infer and reveal sensitive biometric attributes such as race, gender, age, body weight, and eye color; even when such information is not explicitly requested. This raises critical concerns, particularly in real-world applications and socially-sensitive domains. Despite increasing awareness, no publicly available dataset or benchmark exists to comprehensively evaluate or mitigate biometric leakage in MLLMs. To address this gap, we introduce PRISM (Privacy-aware Evaluation of Responses in Sensitive Modalities), a new benchmark designed to assess MLLMs on two fronts: (1) refuse biometric-related queries and (2) implicit biometric leakage in general responses while maintaining semantic faithfulness. Further, we conduct a detailed audit of the widely used LLaVA datasets and uncover extensive biometric leakage across pretraining and instruction data. To address this, we present Safe-LLaVA dataset, the first privacy-preserving MLLM training dataset constructed by systematically removing explicit and implicit biometric information from LLaVA dataset. Our evaluations on PRISM reveal biometric leakages across MLLMs for different attributes, highlighting the detailed privacy-violations. We also fine-tune a model on Safe-LLaVA dataset and show that it substantially reduces the biometric leakages. Together, Safe-LLaVA and PRISM set a new standard for privacy-aligned development and evaluation of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。