arXiv:2604.11668cs.CV2026-04被引 2

统一融合五种地理数据,实现跨模态精准检索与推理

UNIGEOCLIP: Unified Geospatial Contrastive Learning

  • 采用全对全对比学习,统一映射五类地理模态
  • 在多个任务上超越单模态模型和仅用坐标的基线
  • 适合需要多源地理信息融合的科研与应用开发

随着航拍影像、街景图像、高程模型、文本描述及地理坐标等共位地理数据日益丰富,为多模态表征学习提供了独特机遇。本文提出 UNIGEOCLIP,一种大规模多模态对比学习框架,可将五种互补的地理模态统一对齐至单一嵌入空间。不同于以往依赖模态融合或中心枢轴表示的方法,本方法实现所有模态间的全对全对比对齐,支持任意模态组合间的无缝比较、检索与推理。此外,我们设计了扩展的经纬度编码器,通过捕捉多尺度地理结构提升空间表征能力。在多个下游地理任务上的实验表明,UNIGEOCLIP 均显著优于单模态对比模型与仅使用坐标的基线,凸显整体多模态地理对齐的优势。参考实现见 https://gastruc.github.io/unigeoclip。

原文摘要 · Abstract (English)

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.

地理表征多模态学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。