arXiv:2412.06781cs.CVcs.LG2024-12CVPR被引 30

用生成模型让照片定位更精准,还能给出可能位置的概率分布。

Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

  • 基于扩散模型和黎曼流匹配,直接在地球表面去噪定位。
  • 在三个数据集上达到当前最优,且能输出位置概率分布。
  • 适合需要不确定性评估的地理定位场景,如遥感或导航。

全球视觉地理定位旨在预测图像拍摄的具体位置。由于图像的定位精度差异显著,该任务天然存在高度不确定性。然而,现有方法均为确定性模型,忽略了这一特性。本文首次提出基于扩散模型与黎曼流匹配的生成式地理定位方法,其去噪过程直接作用于地球表面。模型在三个基准数据集——OpenStreetView-5M、YFCC-100M 和 iNat21 上均取得领先性能。此外,我们引入概率化视觉地理定位任务,即模型输出所有可能位置的概率分布,而非单一坐标点,并设计了新指标与基线验证方法优势。代码与模型将公开。

原文摘要 · Abstract (English)

Global visual geolocation predicts where an image was captured on Earth. Since images vary in how precisely they can be localized, this task inherently involves a significant degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we aim to close the gap between traditional geolocalization and modern generative methods. We propose the first generative geolocation approach based on diffusion and Riemannian flow matching, where the denoising process operates directly on the Earth's surface. Our model achieves state-of-the-art performance on three visual geolocation benchmarks: OpenStreetView-5M, YFCC-100M, and iNat21. In addition, we introduce the task of probabilistic visual geolocation, where the model predicts a probability distribution over all possible locations instead of a single point. We introduce new metrics and baselines for this task, demonstrating the advantages of our diffusion-based approach. Codes and models will be made available.

视觉定位生成模型扩散模型地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。