解决遥感图像缺模态时的分割难题,提升模型鲁棒性。
DIS2: Disentanglement Meets Distillation with Classwise Attention for Robust Remote Sensing Segmentation under Missing Modalities
- 通过解耦与蒸馏协同机制,主动补偿缺失模态信息。
- 在多基准测试中超越现有方法,显著提升分割精度。
- 适合处理异构、尺度差异大的遥感数据场景。
遥感领域多模态学习因模态缺失而效果大打折扣,且数据异质性强、尺度变化大,导致其他领域有效的方法难以适用。传统解耦学习依赖模态间特征重叠,难以应对这种异质性;知识蒸馏则因学生无法聚焦必要补偿知识,导致语义鸿沟未被弥合。为此,本文提出DIS2新范式,基于三个专为遥感设计的支柱:(1)有原则的缺失信息补偿,(2)类别特定的模态贡献建模,(3)多分辨率特征重要性建模。核心创新是重构解耦与蒸馏的协同关系(称DLKD),显式捕获补偿特征,并与可用模态特征融合,逼近全模态理想表示。类特定特征学习模块(CFLM)自适应学习每类目标在信号可用时的判别证据。二者均依托分层混合融合(HF)结构,利用跨分辨率特征增强预测。大量实验证明,该方法在多个基准上显著优于当前最优方法。
原文摘要 · Abstract (English)
The efficacy of multimodal learning in remote sensing (RS) is severely undermined by missing modalities. The challenge is exacerbated by the RS highly heterogeneous data and huge scale variation. Consequently, paradigms proven effective in other domains often fail when confronted with these unique data characteristics. Conventional disentanglement learning, which relies on significant feature overlap between modalities (modality-invariant), is insufficient for this heterogeneity. Similarly, knowledge distillation becomes an ill-posed mimicry task where a student fails to focus on the necessary compensatory knowledge, leaving the semantic gap unaddressed. Our work is therefore built upon three pillars uniquely designed for RS: (1) principled missing information compensation, (2) class-specific modality contribution, and (3) multi-resolution feature importance. We propose a novel method DIS2, a new paradigm shifting from modality-shared feature dependence and untargeted imitation to active, guided missing features compensation. Its core novelty lies in a reformulated synergy between disentanglement learning and knowledge distillation, termed DLKD. Compensatory features are explicitly captured which, when fused with the features of the available modality, approximate the ideal fused representation of the full-modality case. To address the class-specific challenge, our Classwise Feature Learning Module (CFLM) adaptively learn discriminative evidence for each target depending on signal availability. Both DLKD and CFLM are supported by a hierarchical hybrid fusion (HF) structure using features across resolutions to strengthen prediction. Extensive experiments validate that our proposed approach significantly outperforms state-of-the-art methods across benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。