构建250万样本遥感多模态数据集,实现语义对齐的模型预训练
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing

- 统一对齐协议构建跨模态遥感数据,支持多分辨率融合
- 通过智能生成与验证框架,实现带语义标注的文本描述
- 适用于遥感图像理解、跨传感器泛化等下游任务
遥感领域高效基础建模需要空间对齐的异构模态与语义接地的监督,但此类资源规模有限。我们提出GeoMeld,一个包含约250万组空间对齐样本的大规模多模态数据集,覆盖多种模态和分辨率,并采用统一对齐协议支持模态感知表征学习。该数据集通过代理式标题生成框架,结合光谱信号、地形统计量和结构化地理元数据,合成并验证注释,使文本描述编码可度量的跨模态关系。为利用该数据集,我们引入GeoMeld-FM预训练框架,整合多前缀掩码自编码、JEPA表征学习与图文对比对齐,联合目标使表征空间同时捕捉可靠跨传感器物理一致性与语义接地性。实验表明在下游迁移与跨传感器鲁棒性上持续提升。GeoMeld与GeoMeld-FM共同建立遥感多模态基础建模的可扩展基准框架。
原文摘要 · Abstract (English)
Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-scale multimodal dataset with approximately 2.5 million spatially aligned samples. The dataset spans diverse modalities and resolutions and is constructed under a unified alignment protocol for modality-aware representation learning. GeoMeld provides semantically grounded language supervision through an agentic captioning framework that synthesizes and verifies annotations from spectral signals, terrain statistics, and structured geographic metadata, encoding measurable cross-modality relationships within textual descriptions. To leverage this dataset, we introduce GeoMeld-FM, a pretraining framework that combines multi-pretext masked autoencoding over aligned modalities, JEPA representation learning, and caption-vision contrastive alignment. This joint objective enables the learned representation space to capture both reliable cross-sensor physical consistency and grounded semantics. Experiments demonstrate consistent gains in downstream transfer and cross-sensor robustness. Together, GeoMeld and GeoMeld-FM establish a scalable reference framework for semantically grounded multi-modal foundation modeling in remote sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。