arXiv:2606.20167cs.LG2026-06被引 1

用多模态对比学习构建隐式地理嵌入,提升空间预测性能

Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying

  • 通过位置绑定设计多模态对比学习框架,利用未配对地理数据
  • 在四个下游任务中达到最强双模态基线表现,但模态越多越不增效
  • MELT训练更稳定,适合未来扩展,揭示编码器是主要瓶颈

空间预测任务常受限于高质量标注数据的缺乏。自监督预训练,特别是对比学习,是潜在解决方案,但现有方法通常仅将地理坐标与单一模态对齐。本文提出两种多模态对比学习架构:基于位置绑定的多模态嵌入(MELT)与顺序交替位置训练(SALT),拓展了该框架至多于两个模态,利用未配对的地理空间数据。两种方法技术上可行,在四个下游任务中均达到最强双模态基线(SATCLIP)的性能。然而,增加模态数量并未持续提升效果,表明位置编码器是主要限制——对比目标在早期即达上限,与模态多样性或预训练量无关。MELT相比SALT训练更稳定,为未来扩展提供更强基础。

原文摘要 · Abstract (English)

Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solution, with contrastive learning dominant for location encoders. Those approaches usually align geographic coordinates with just one additional modality. We propose two multimodal contrastive learning architectures: Multimodal Embedding via Location Tying (MELT) and Sequential Alternating Location Training (SALT). These architectures expand this framework beyond two modalities by utilising unpaired geospatial data. Both methods are technically viable and match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks. However, increasing the number of modalities does not consistently improve performance, suggesting that the chosen location encoder is the main limitation - the contrastive objective reaches its peak early, regardless of modality diversity or pre-training volume. MELT provides more stable training than SALT and presents a stronger foundation for future scaling.

多模态学习对比学习地理嵌入自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。