评估生成式视觉语言模型的定位能力与隐私风险
Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models
- 测试25个顶尖视觉语言模型在四大数据集上的定位能力
- 社交媒体风格图像定位准确率达61%,普通街景图表现差
- 揭示模型潜在隐私泄露风险,提醒公众警惕照片外泄
地理定位是指仅通过视觉线索识别图像位置的任务,具有提升灾后响应、导航和地理教育等应用价值。近年来,视觉语言模型(VLMs)展现出高精度图像定位能力,但这也带来严重隐私风险,如跟踪与监控,尤其在社交媒体广泛传播照片的背景下。当前对生成式VLMs的定位精度、局限性及无意推断潜力缺乏系统评估。为此,我们对25个前沿VLMs在四个涵盖多样化环境的基准图像数据集上进行了全面评估。结果揭示了VLMs内部推理机制,凸显其优势、局限及社会风险:模型在通用街景图像上表现不佳,但在类似社交媒体内容的图像上达到61%的高准确率,引发紧迫的隐私担忧。
原文摘要 · Abstract (English)
Geo-localization is the task of identifying the location of an image using visual cues alone. It has beneficial applications, such as improving disaster response, enhancing navigation, and geography education. Recently, Vision-Language Models (VLMs) are increasingly demonstrating capabilities as accurate image geo-locators. This brings significant privacy risks, including those related to stalking and surveillance, considering the widespread uses of AI models and sharing of photos on social media. The precision of these models is likely to improve in the future. Despite these risks, there is little work on systematically evaluating the geolocation precision of Generative VLMs, their limits and potential for unintended inferences. To bridge this gap, we conduct a comprehensive assessment of the geolocation capabilities of 25 state-of-the-art VLMs on four benchmark image datasets captured in diverse environments. Our results offer insight into the internal reasoning of VLMs and highlight their strengths, limitations, and potential societal risks. Our findings indicate that current VLMs perform poorly on generic street-level images yet achieve notably high accuracy (61\%) on images resembling social media content, raising significant and urgent privacy concerns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。