arXiv:2505.19139cs.CV2025-05被引 15

用视觉语言模型从个人照片推断隐私属性,发现新隐私风险

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework

论文配图:The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
图 1 · 摘自论文原文
  • 构建混合智能体框架HolmesEye,结合视觉与语言模型分析图像内外关系
  • 在251人3012项属性数据上,抽象属性预测超越人类15%准确率
  • 首个大规模多图隐私属性数据集PAPI,助力未来隐私研究

本研究揭示了视觉语言模型(VLM)智能体框架带来的新型隐私风险:仅凭一组个人照片即可推断敏感属性(如年龄、健康状况)甚至抽象属性(如性格、社交特征),我们称之为“图像隐私属性画像”。该风险尤为严重,因现代应用可轻易访问用户相册,且图像集合推理能利用图像间关联实现更精细的画像。然而,当前面临两大挑战:缺乏带多图标注的隐私属性基准数据集,以及现有多模态大模型在大量图像中推断抽象属性能力有限。为此,我们构建了目前最大的个人图像隐私属性研究数据集PAPI,包含251名个体的2,510张图像及3,012个标注隐私属性。同时提出HolmesEye混合智能体框架,结合VLM提取图像内与图像间信息,借助LLM引导推理并进行取证式结果整合,突破长上下文视觉推理瓶颈。实验表明,HolmesEye在平均准确率上比顶尖基线提升10.8%,在抽象属性预测上超越人类水平15.0%。该工作凸显图像画像隐私风险的紧迫性,并提供了新数据集与先进框架以推动后续研究。

原文摘要 · Abstract (English)

Our research reveals a new privacy risk associated with the vision-language model (VLM) agentic framework: the ability to infer sensitive attributes (e.g., age and health information) and even abstract ones (e.g., personality and social traits) from a set of personal images, which we term "image private attribute profiling." This threat is particularly severe given that modern apps can easily access users' photo albums, and inference from image sets enables models to exploit inter-image relations for more sophisticated profiling. However, two main challenges hinder our understanding of how well VLMs can profile an individual from a few personal photos: (1) the lack of benchmark datasets with multi-image annotations for private attributes, and (2) the limited ability of current multimodal large language models (MLLMs) to infer abstract attributes from large image collections. In this work, we construct PAPI, the largest dataset for studying private attribute profiling in personal images, comprising 2,510 images from 251 individuals with 3,012 annotated privacy attributes. We also propose HolmesEye, a hybrid agentic framework that combines VLMs and LLMs to enhance privacy inference. HolmesEye uses VLMs to extract both intra-image and inter-image information and LLMs to guide the inference process as well as consolidate the results through forensic analysis, overcoming existing limitations in long-context visual reasoning. Experiments reveal that HolmesEye achieves a 10.8% improvement in average accuracy over state-of-the-art baselines and surpasses human-level performance by 15.0% in predicting abstract attributes. This work highlights the urgency of addressing privacy risks in image-based profiling and offers both a new dataset and an advanced framework to guide future research in this area.

隐私安全视觉语言模型属性推断数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。