arXiv:2609.03391cs.CVcs.AI2026-09

让CLIP模型读懂雷达、多光谱等异源遥感数据

Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data

论文配图:Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data
图 1 · 摘自论文原文
  • 通过基分解技术实现多源传感器输入的统一表征
  • 在500万规模数据集上实现跨模态检索与零样本分类
  • 适合遥感多模态理解、跨域图像分析的研究者

对比语言-图像学习(CLIP)已成为遥感视觉-语言理解的关键范式。然而,现有方法多基于面向RGB的CLIP架构,难以利用合成孔径雷达(SAR)、多光谱成像(MSI)和高光谱成像(HSI)等异源传感器。为此,我们提出OmniRSCLIP,一种支持多源传感器输入的端到端对比学习框架。核心思想是将固定通道输入扩展为可适配任意通道,同时不破坏预训练视觉知识。OmniRSCLIP引入光谱-空间基分解(SSBD),将任意通道适配建模为基重构问题:预训练的CLIP块嵌入提供可迁移的空间基,波长条件系数在受限视觉先验空间内构建传感器特异性嵌入核。该设计避免强制异源传感器进入固定通道空间,同时将其对齐至统一的图文语义空间。此外,我们提出一种光谱-上下文感知的掩码对比学习方案,抑制模态特异性冗余特征,增强细粒度图文对齐。最后,为支持多模态训练,我们构建了首个涵盖RGB、SAR、MSI和HSI的大规模遥感图文语料库OmniRS5M。在检索、零样本分类和语义定位任务上的实验表明,OmniRSCLIP在保持强RGB域性能的同时,有效拓展了CLIP对异源遥感模态的支持能力。

原文摘要 · Abstract (English)

Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.

遥感多模态对比学习视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。