提出统一多模态框架M²CD,提升光学与SAR图像变化检测性能
M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation
- 引入专家混合模块显式处理多模态差异,增强跨模态特征学习
- 设计光学到SAR引导路径并结合自蒸馏,显著缩小模态间特征差距
- 在多种骨干网络上验证有效,最优模型性能超越现有最好方法
现有变化检测(CD)方法多聚焦于不同时相的光学图像,深度学习在此领域已取得显著进展。但在灾害响应等极端场景中,具有主动成像能力的合成孔径雷达(SAR)更适合作为灾后数据源。这给CD方法带来新挑战:现有共享权重的孪生网络难以有效学习光学与SAR图像间的跨模态分布。为此,我们提出统一多模态变化检测框架M²CD。通过在主干网络中集成专家混合(MoE)模块,显式处理不同模态特性,增强模型对多模态数据分布的学习能力。同时,创新性地提出光学到SAR引导路径(O2SP),并在训练中引入自蒸馏机制,进一步缩小不同模态间的特征空间差异,减轻模型学习负担。基于CNN与Transformer骨干网络设计了多个M²CD变体。大量实验验证框架有效性,其中采用MiT-b1骨干的M²CD版本在光学-SAR变化检测任务中优于所有现有最先进方法。
原文摘要 · Abstract (English)
Most existing change detection (CD) methods focus on optical images captured at different times, and deep learning (DL) has achieved remarkable success in this domain. However, in extreme scenarios such as disaster response, synthetic aperture radar (SAR), with its active imaging capability, is more suitable for providing post-event data. This introduces new challenges for CD methods, as existing weight-sharing Siamese networks struggle to effectively learn the cross-modal data distribution between optical and SAR images. To address this challenge, we propose a unified MultiModal CD framework, M$^2$CD. We integrate Mixture of Experts (MoE) modules into the backbone to explicitly handle diverse modalities, thereby enhancing the model's ability to learn multimodal data distributions. Additionally, we innovatively propose an Optical-to-SAR guided path (O2SP) and implement self-distillation during training to reduce the feature space discrepancy between different modalities, further alleviating the model's learning burden. We design multiple variants of M$^2$CD based on both CNN and Transformer backbones. Extensive experiments validate the effectiveness of the proposed framework, with the MiT-b1 version of M$^2$CD outperforming all state-of-the-art (SOTA) methods in optical-SAR CD tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。