arXiv:2503.16683cs.CVcs.AI2025-03中稿 · ISPRS Journal of P…被引 16

提出GAIR模型,让遥感图像在任意位置都有精准定位表示。

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations

  • 用隐式神经表示扩展ViT,实现任意位置的连续遥感图像表征
  • 融合遥感、街景和地理信息,通过对比学习训练无监督模型
  • 在9个任务22个数据集上超越现有方法,适合多模态地理空间分析

视觉变换器(ViT)在计算机视觉任务中表现优异,但缺乏在地理空间任务中对任意位置的精细局部表征。此类任务涉及高分辨率遥感(RS)数据、地面街景(SV)影像及地理矢量数据,需高分辨率局部表示以建模跨模态地理关系与对齐。本文提出一种基于隐式神经表示(INR)模块的扩展结构,支持任意位置的连续遥感图像表征。在此基础上,构建了新的自监督学习框架GAIR,融合遥感、街景与地理位置元数据,使用三个因子化神经编码器将不同模态投影至共享嵌入空间,并通过INR模块进行地理对齐。模型在无标签数据上采用对比学习目标进行训练。在9个地理空间任务和22个数据集上评估,涵盖基于遥感图像、街景图像和位置嵌入的基准测试。实验表明,GAIR优于现有地理基础模型(GeoFM)及未使用细粒度地理对齐空间表示的对比学习方法(如MoCo V3和MAE)。结果验证了其在跨任务、多尺度与多时相场景下学习通用地理空间表征的有效性。项目代码已开源:https://github.com/zpl99/GAIR。

原文摘要 · Abstract (English)

Vision Transformer (ViT) has been widely used in computer vision tasks with excellent results by providing representations for a whole image or image patches. However, ViT lacks detailed localized image representations at arbitrary positions when applied to geospatial tasks that involve multiple geospatial data modalities, such as overhead remote sensing (RS) data, ground-level imagery, and geospatial vector data. Here high-resolution localized representations are vital for modeling geospatial relationships and alignments across modalities. We proposed to solve this representation problem with an implicit neural representation (INR) module extending ViT with Neural Implicit Local Interpolation, which produces a continuous RS image representation covering arbitrary location in the RS image. Based on the INR module, we introduce GAIR, a novel location-aware self-supervised learning (SSL) objective integrating overhead RS data, street view (SV) imagery, and their geolocation metadata. GAIR utilizes three factorized neural encoders to project different modalities into the embedding space, and the INR module is used to further align these representations geographically, which are trained with contrastive learning objectives from unlabeled data. We evaluate GAIR across 9 geospatial tasks and 22 datasets spanning RS image-based, SV image-based, and location embedding-based benchmarks. Experimental results demonstrate that GAIR outperforms state-of-the-art geo-foundation models (GeoFM) and alternative SSL training objectives (e.g., MoCo V3 and MAE) that do not use fine-grained geo-aligned spatial representations. Our results highlight the effectiveness of GAIR in learning generalizable geospatial representations across tasks, spatial scales, and temporal contexts. The project code is available at https://github.com/zpl99/GAIR.

地理表征自监督学习隐式表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。