arXiv:2503.18142cs.CVcs.AI2025-03NeurIPS被引 7

用扩散模型在球谐空间定位照片拍摄地,提升跨区域泛化能力。

LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space

  • 构建球谐狄拉克编码框架,将地理坐标转为多尺度希尔伯特空间表示
  • 在5个全球数据集上超越现有方法,对未见地点仍具强泛化性
  • 首个在潜空间进行多尺度扩散定位的模型,适合跨区域图像定位任务

图像地理定位是推断照片拍摄地的重要但具挑战性的任务。现有方法依赖网格分类或图库检索,当测试图像空间分布与网格或图库不匹配时,泛化能力显著下降。新兴生成式方法虽摆脱网格和图库,但使用原始地理坐标,因缺乏多尺度信息导致定位质量下降。为此,我们提出多尺度潜在扩散模型LocDiff。设计新型球谐狄拉克编码解码框架(SHDD),将地球表面点编码为球谐系数的希尔伯特空间表示,并通过球面概率分布的模式搜索解码位置。提出基于SirenNet的架构CS-UNet,通过最小化潜在KL散度损失,在潜空间学习图像条件下的反向扩散过程。据我们所知,LocDiff是首个在多尺度位置编码空间中执行潜在扩散并由图像引导生成地理坐标的图像地理定位模型。实验表明,该模型在5个具有挑战性的全球尺度图像地理定位数据集上均优于所有先进基线方法,且对未见地理区域表现出显著更强的泛化能力。

原文摘要 · Abstract (English)

Image geolocalization is a fundamental yet challenging task, aiming at inferring the geolocation on Earth where an image is taken. State-of-the-art methods employ either grid-based classification or gallery-based image-location retrieval, whose spatial generalizability significantly suffers if the spatial distribution of test images does not align with the choices of grids and galleries. Recently emerging generative approaches, while getting rid of grids and galleries, use raw geographical coordinates and suffer quality losses due to their lack of multi-scale information. To address these limitations, we propose a multi-scale latent diffusion model called LocDiff for image geolocalization. We developed a novel positional encoding-decoding framework called Spherical Harmonics Dirac Delta (SHDD) Representations, which encodes points on a spherical surface (e.g., geolocations on Earth) into a Hilbert space of Spherical Harmonics coefficients and decodes points (geolocations) by mode-seeking on spherical probability distributions. We also propose a novel SirenNet-based architecture (CS-UNet) to learn an image-based conditional backward process in the latent SHDD space by minimizing a latent KL-divergence loss. To the best of our knowledge, LocDiff is the first image geolocalization model that performs latent diffusion in a multi-scale location encoding space and generates geolocations under the guidance of images. Experimental results show that LocDiff can outperform all state-of-the-art grid-based, retrieval-based, and diffusion-based baselines across 5 challenging global-scale image geolocalization datasets, and demonstrates significantly stronger generalizability to unseen geolocations.

图像定位扩散模型球谐编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。