arXiv:2503.06312cs.CV2025-03被引 9

统一多模态遥感视觉语言模型,支持多种传感器动态适配。

DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation

  • 用单一Transformer架构动态适应六类遥感数据
  • 零样本测试在多种未见波段和模态上达顶尖性能
  • 适合遥感、地理信息与大模型融合研究者

地球观测(EO)涵盖光学、雷达、多光谱和高光谱等多种模态,各自捕捉不同的环境信号。然而,当前基于CLIP的遥感视觉语言模型仍局限于单一模态,限制了跨任务的泛化与可扩展性。我们提出DOFA-CLIP(Dynamic-One-For-All CLIP),一种统一的视觉语言基础模型,通过单个Transformer主干网络动态适应不同遥感模态及灵活光谱配置。主要贡献包括:1)构建GeoLangBind-2M,一个覆盖六类异构模态的大规模遥感图文数据集,含丰富自然语言描述;2)提出VECT(Vision-models Enhanced Contrastive Text-image pretraining)训练策略,利用多个视觉基础模型增强CLIP特征的空间感知能力;3)设计模态感知知识聚合(MaKA)模块,实现具有模态特异性感知的特征蒸馏。DOFA-CLIP在多种遥感基准上实现领先零样本性能,涵盖未见过的模态和不同数量的输入光谱波段。这些成果为多模态遥感理解建立可扩展基础,并推动异构遥感数据与大语言模型的融合。代码与数据集将公开发布。

原文摘要 · Abstract (English)

Earth observation (EO) spans a broad spectrum of modalities, including optical, radar, multispectral, and hyperspectral data, each capturing distinct environmental signals. However, current vision-language models in EO, particularly CLIP-based variants, remain confined to individual modalities, limiting generalization and scalability across diverse tasks. We present DOFA-CLIP (Dynamic-One-For-All CLIP), a unified vision-language foundation model that dynamically adapts to EO modalities with flexible spectral configurations through a single Transformer backbone. Our approach introduces three key contributions: 1) the construction of GeoLangBind-2M, a large-scale EO image-text dataset covering six heterogeneous modalities with rich natural language descriptions; 2) a novel training strategy called VECT (Vision-models Enhanced Contrastive Text-image pretraining), which enhances the spatial awareness of CLIP features with multiple vision foundation models; and 3) a Modality-aware Knowledge Agglomeration (MaKA) module that refines feature distillation with modality-specific awareness. DOFA-CLIP achieves state-of-the-art zero-shot performance across a wide range of EO benchmarks, including unseen modalities and a diverse number of input spectral bands. Together, these contributions establish a scalable foundation for multimodal EO understanding and open new avenues for integrating heterogeneous EO data with large language models. Code and datasets will be released. Code and datasets are publicly available.

遥感多模态大模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。