arXiv:2507.20259cs.CV2025-07

用少量标签实现卫星图像高效分类,无需像素级配对数据。

L-MCAT: Unpaired Multimodal Transformer with Contrastive Attention for Label-Efficient Satellite Image Classification

  • 通过模态-光谱适配器压缩多源遥感数据,统一嵌入空间。
  • 无配对注意力对齐机制在仅20标签/类下达到95.4%准确率。
  • 轻量设计适合单卡训练,对空间错位有强鲁棒性,适合实际应用。

我们提出轻量级多模态对比注意力变换器(L-MCAT),一种基于Transformer的框架,用于利用未配对的多模态卫星数据实现标签高效的遥感图像分类。L-MCAT引入两项核心创新:(1) 模态-光谱适配器(MSA),将高维传感器输入压缩至统一嵌入空间;(2) 无配对多模态注意力对齐(U-MAA),一种嵌入注意力层的对比自监督机制,可在无像素级对应或标签的情况下对齐异构模态。L-MCAT在SEN12MS数据集上仅使用每类20个标签即达到95.4%总体准确率,优于现有最优基线,且参数量仅为MCTrans的1/47,浮点运算量为1/23。即使在50%空间错位下仍保持超92%准确率,展现出强鲁棒性。模型可在单张消费级显卡上于5小时内完成端到端训练。

原文摘要 · Abstract (English)

We propose the Lightweight Multimodal Contrastive Attention Transformer (L-MCAT), a novel transformer-based framework for label-efficient remote sensing image classification using unpaired multimodal satellite data. L-MCAT introduces two core innovations: (1) Modality-Spectral Adapters (MSA) that compress high-dimensional sensor inputs into a unified embedding space, and (2) Unpaired Multimodal Attention Alignment (U-MAA), a contrastive self-supervised mechanism integrated into the attention layers to align heterogeneous modalities without pixel-level correspondence or labels. L-MCAT achieves 95.4% overall accuracy on the SEN12MS dataset using only 20 labels per class, outperforming state-of-the-art baselines while using 47x fewer parameters and 23x fewer FLOPs than MCTrans. It maintains over 92% accuracy even under 50% spatial misalignment, demonstrating robustness for real-world deployment. The model trains end-to-end in under 5 hours on a single consumer GPU.

遥感分类多模态少样本轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。