大模型能从图片精准定位地理位置,引发隐私安全新风险。
Evaluation of Geolocation Capabilities of Multimodal Large Language Models and Analysis of Associated Privacy Risks
- 用视觉推理模型分析街景图的地理线索。
- 顶尖模型在1公里范围内定位准确率达49%。
- 适合关注AI隐私与安全的研究者阅读。
多模态大模型(MLLMs)的快速发展显著提升了其推理能力,广泛应用于各类智能场景。然而,这些进步也带来了严重的隐私与伦理问题。如今,仅凭图像内容,MLLMs即可推断社交媒体或街景图片的地理位置,可能引发泄露身份、监控等安全威胁。本研究系统梳理了基于MLLM的地理定位技术文献,评估了前沿视觉推理模型在街景图像定位任务中的表现。实证结果表明,最先进模型对街景图像的定位准确率可达49%(1公里半径内),显示出其从视觉数据中提取精细地理线索的强大能力。研究进一步识别出影响定位成功的关键视觉要素,如文字、建筑风格和环境特征,并探讨了相关隐私风险及技术与政策层面的应对措施。代码与数据集已开源:https://github.com/zxyl1003/MLLM-Geolocation-Evaluation。
原文摘要 · Abstract (English)
Objectives: The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly enhanced their reasoning capabilities, enabling a wide range of intelligent applications. However, these advancements also raise critical concerns regarding privacy and ethics. MLLMs are now capable of inferring the geographic location of images -- such as those shared on social media or captured from street views -- based solely on visual content, thereby posing serious risks of privacy invasion, including doxxing, surveillance, and other security threats. Methods: This study provides a comprehensive analysis of existing geolocation techniques based on MLLMs. It systematically reviews relevant litera-ture and evaluates the performance of state-of-the-art visual reasoning models on geolocation tasks, particularly in identifying the origins of street view imagery. Results: Empirical evaluation reveals that the most advanced visual large models can successfully localize the origin of street-level imagery with up to $49\%$ accuracy within a 1-kilometer radius. This performance underscores the models' powerful capacity to extract and utilize fine-grained geographic cues from visual data. Conclusions: Building on these findings, the study identifies key visual elements that contribute to suc-cessful geolocation, such as text, architectural styles, and environmental features. Furthermore, it discusses the potential privacy implications associated with MLLM-enabled geolocation and discuss several technical and policy-based coun-termeasures to mitigate associated risks. Our code and dataset are available at https://github.com/zxyl1003/MLLM-Geolocation-Evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。