解决遥感多模态数据缺失问题,提升分割精度
STARS: Shared-specific Translation and Alignment for missing-modality Remote Sensing Semantic Segmentation
- 设计双向翻译与梯度截断的非对称对齐机制
- 像素级语义采样对齐策略提升小类识别能力
- 适用于光学或数字高程图缺失的实用场景
多模态遥感技术通过融合光学图像、合成孔径雷达(SAR)和数字表面模型(DSM)等异构数据,显著提升了地表语义理解能力。然而在实际应用中,模态数据缺失(如缺少光学图像或DSM)是常见且严重的问题,导致传统多模态融合模型性能下降。现有方法仍存在特征坍塌和恢复特征过于泛化等问题。为此,我们提出STARS(Shared-specific Translation and Alignment for missing-modality Remote Sensing),一种针对不完整多模态输入的鲁棒语义分割框架。STARS基于两项核心设计:首先,引入带有双向翻译和梯度截断的非对称对齐机制,有效防止特征坍塌并降低对超参数的敏感性;其次,提出像素级语义采样对齐(PSA)策略,结合类别平衡的像素采样与跨模态语义对齐损失,缓解因严重类别不平衡导致的对齐失败,提升少数类别识别能力。
原文摘要 · Abstract (English)
Multimodal remote sensing technology significantly enhances the understanding of surface semantics by integrating heterogeneous data such as optical images, Synthetic Aperture Radar (SAR), and Digital Surface Models (DSM). However, in practical applications, the missing of modality data (e.g., optical or DSM) is a common and severe challenge, which leads to performance decline in traditional multimodal fusion models. Existing methods for addressing missing modalities still face limitations, including feature collapse and overly generalized recovered features. To address these issues, we propose \textbf{STARS} (\textbf{S}hared-specific \textbf{T}ranslation and \textbf{A}lignment for missing-modality \textbf{R}emote \textbf{S}ensing), a robust semantic segmentation framework for incomplete multimodal inputs. STARS is built on two key designs. First, we introduce an asymmetric alignment mechanism with bidirectional translation and stop-gradient, which effectively prevents feature collapse and reduces sensitivity to hyperparameters. Second, we propose a Pixel-level Semantic sampling Alignment (PSA) strategy that combines class-balanced pixel sampling with cross-modality semantic alignment loss, to mitigate alignment failures caused by severe class imbalance and improve minority-class recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。