arXiv:2503.05774cs.LGcs.DB2025-03被引 2

GeoJEPA消除地理多模态学习中的数据增强与采样偏差,提升模型泛化能力。

GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

  • 基于联合嵌入预测架构,无需依赖人工设计的预训练任务
  • 在OpenStreetMap大规模数据上自监督预训练,生成城市区域语义表示
  • 适用于城市规划、地图理解等需要多模态融合的地理研究场景

现有地理空间区域与地图实体的自监督表征学习方法严重依赖预训练任务设计,常通过空间邻近性进行数据增强或正负样本采样,导致引入偏差并限制表征能力与泛化性。为此,本文提出GeoJEPA,一种基于自监督联合嵌入预测架构的通用多模态融合模型,旨在消除自监督地理表征学习中广泛存在的增强与采样偏差。GeoJEPA在包含OpenStreetMap属性、几何信息和航拍图像的大规模数据集上进行自监督预训练,生成城市区域与地图实体的多模态语义表示,并通过定量与定性评估验证其性能。本工作揭示了JEPA在处理多模态地理数据方面的关键优势。

原文摘要 · Abstract (English)

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and negative pairs based on spatial proximity. This reliance introduces biases and limits the representations' expressiveness and generalisability. Consequently, the literature has expressed a pressing need to explore different methods for modelling geospatial data. To address the key difficulties of such methods, namely multimodality, heterogeneity, and the choice of pretext tasks, we present GeoJEPA, a versatile multimodal fusion model for geospatial data built on the self-supervised Joint-Embedding Predictive Architecture. With GeoJEPA, we aim to eliminate the widely accepted augmentation- and sampling biases found in self-supervised geospatial representation learning. GeoJEPA uses self-supervised pretraining on a large dataset of OpenStreetMap attributes, geometries and aerial images. The results are multimodal semantic representations of urban regions and map entities that we evaluate both quantitatively and qualitatively. Through this work, we uncover several key insights into JEPA's ability to handle multimodal data.

地理表征多模态学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。