arXiv:2507.08683cs.CVcs.AI2025-07被引 3

多模态遥感图像中,用标签感知对比学习提升低标注场景下的分类精度。

MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing

  • 融合跨模态与模内对比学习,引入多标签监督增强语义对齐
  • 在低标注、高类别重叠下准确率超越自监督与全监督基线
  • 适合处理光学与雷达图像融合的复杂遥感场景分析

对比学习(CL)作为无须大量标注数据即可学习可迁移表征的强大范式,在计算机视觉任务中表现优异。其捕捉样本间内在相似性与差异的能力使其特别适用于地球系统观测(ESO),因为光学与合成孔径雷达(SAR)影像能自然对齐同一地理区域。然而,ESO面临高类别相似性、场景杂乱和边界模糊等挑战,尤其在低标注、多标签设置下更难学习有效表征。现有对比学习框架通常只关注模内自监督,缺乏跨模态多标签对齐与语义精确性机制。本文提出MoSAiC框架,统一优化模内与跨模态对比学习,并引入多标签监督对比损失。该框架专为多模态卫星影像设计,实现更精细的语义解耦与鲁棒表征学习,尤其在光谱相似与空间复杂的类别中表现突出。在两个基准数据集BigEarthNet V2.0与Sent12MS上的实验表明,MoSAiC在低标注及高类别重叠场景下,于准确率、聚类一致性与泛化能力方面均持续优于全监督与自监督基线。

原文摘要 · Abstract (English)

Contrastive learning (CL) has emerged as a powerful paradigm for learning transferable representations without the reliance on large labeled datasets. Its ability to capture intrinsic similarities and differences among data samples has led to state-of-the-art results in computer vision tasks. These strengths make CL particularly well-suited for Earth System Observation (ESO), where diverse satellite modalities such as optical and SAR imagery offer naturally aligned views of the same geospatial regions. However, ESO presents unique challenges, including high inter-class similarity, scene clutter, and ambiguous boundaries, which complicate representation learning -- especially in low-label, multi-label settings. Existing CL frameworks often focus on intra-modality self-supervision or lack mechanisms for multi-label alignment and semantic precision across modalities. In this work, we introduce MoSAiC, a unified framework that jointly optimizes intra- and inter-modality contrastive learning with a multi-label supervised contrastive loss. Designed specifically for multi-modal satellite imagery, MoSAiC enables finer semantic disentanglement and more robust representation learning across spectrally similar and spatially complex classes. Experiments on two benchmark datasets, BigEarthNet V2.0 and Sent12MS, show that MoSAiC consistently outperforms both fully supervised and self-supervised baselines in terms of accuracy, cluster coherence, and generalization in low-label and high-class-overlap scenarios.

遥感图像对比学习多模态多标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。