arXiv:2601.08420cs.CV2026-01中稿 · InGARSS 2025

用语言模型对齐遥感数据与语义,提升多模态理解能力。

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

  • 通过文本嵌入引导,将高光谱与激光雷达数据映射到共享语义空间。
  • 在两个基准数据集上表现优于多个视觉仅方法,证明语言监督有效。
  • 仅用简单CNN编码器即可达到先进水平,适合遥感语义分析场景。

本文提出一种新型多模态框架——多模态语言引导网络(MMLGNet),利用视觉-语言模型(如CLIP)将高光谱成像(HSI)和激光雷达(LiDAR)等异构遥感模态与自然语言语义对齐。随着多源地球观测数据日益丰富,亟需融合光谱、空间与几何信息并实现语义级理解的方法。MMLGNet采用模态特定编码器,通过双向对比学习将视觉特征与手工设计的文本嵌入在共享潜在空间中对齐。受CLIP训练范式启发,该方法弥合了高维遥感数据与语言引导解释之间的鸿沟。值得注意的是,MMLGNet仅使用简单的CNN编码器即在两个基准数据集上超越多个成熟的纯视觉多模态方法,验证了语言监督的显著优势。代码已公开于https://github.com/AdityaChaudhary2913/CLIP_HSI。

原文摘要 · Abstract (English)

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.

遥感多模态CLIP语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。