arXiv:2606.12595cs.LGcs.AI2026-06

对比多种地理空间多模态模型,揭示灵活性与性能的权衡。

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

论文配图:Emerging Flexible Designs for Geospatial Multimodal Foundation Models
图 1 · 摘自论文原文
  • 统一预训练目标和数据集,公平比较不同架构
  • 在GEOBench上验证分类与分割任务表现差异
  • 为构建更鲁棒的地理多模态模型提供设计指导

基础模型正在迅速改变地球观测,通过在多样化的未标注地理空间模态上实现可扩展的预训练。然而,从仅编码器到编码器-解码器及掩码自编码等不同架构范式,使得性能权衡难以一致评估。本文对面向地理空间多模态推理的领先基础模型架构进行严格对比,重点关注其在不同光谱波段配置下的灵活性。所有模型采用相同的自监督学习目标和训练数据集,并在参数量一致条件下,在GEOBench基准上评估分类与分割任务表现。结果揭示了模型灵活性、模态对齐性与下游任务性能之间的设计权衡。在受控条件下明确各架构的优势与局限,为构建下一代具备强大多模态推理能力的地理空间基础模型提供实用指导。

原文摘要 · Abstract (English)

Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities. However, their architectural diversity ranging from encoder-only to encoder-decoder and masked autoencoding paradigms makes it challenging to assess performance trade offs in a consistent manner. In this work, we present an apples-to-apples comparison of leading FM architectures designed for geospatial multimodal reasoning, with a particular focus on flexibility across varied spectral band configurations. We standardize pretraining using identical self supervised learning objectives and training datasets, and evaluate all models under consistent parameterization on the GEOBench benchmark across classification and segmentation tasks. Our results offer new insights into the design trade-offs between model flexibility, modality alignment, and downstream task performance. By highlighting architectural strengths and limitations under controlled conditions, this study provides practical guidance for building next generation geospatial foundation models capable of robust multimodal reasoning.

地理空间多模态基础模型架构对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。