无需专门训练,大模型也能精准定位图片拍摄地。
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
- 用谷歌街景数据构建新基准,评估模型从单图推地理坐标能力。
- 多数基础模型定位误差中位数小于300公里,表现惊人。
- 接入工具的智能体可降低30.6%误差,适合隐私安全研究者关注。
视觉语言模型(VLM)的普及引发了视觉信息时代下的隐私担忧。尽管基础模型具备广泛知识与泛化能力,本文聚焦其对未见过图像的地理定位推理能力。提出一个基于谷歌街景的基准数据集,涵盖全球覆盖分布。评估基础模型在单图地理定位任务中的表现,结果显示多数模型的中位距离误差低于300公里。进一步测试拥有辅助工具的VLM智能体,发现其距离误差最多可降低30.6%。结果表明,现代基础VLM无需专门训练即可作为强大的图像定位工具。结合模型日益开放的可访问性,该能力对在线隐私构成更大威胁。论文讨论相关风险及未来研究方向。
原文摘要 · Abstract (English)
The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabilities, we specifically investigate their ability to infer geographic location from previously unseen image data. This paper introduces a benchmark dataset collected from Google Street View that represents its global distribution of coverage. Foundation models are evaluated on single-image geolocation inference, with many achieving median distance errors of <300 km. We further evaluate VLM "agents" with access to supplemental tools, observing up to a 30.6% decrease in distance error. Our findings establish that modern foundation VLMs can act as powerful image geolocation tools, without being specifically trained for this task. When coupled with increasing accessibility of these models, our findings have greater implications for online privacy. We discuss these risks, as well as future work in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。