让分割模型不依赖特定传感器,任意组合都能稳准运行。
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
- 用多模态并行学习训练强教师模型,再通过特征级蒸馏传递知识。
- 在多尺度空间实现跨模态与单模态知识迁移,解决模态依赖问题。
- 适合真实场景中传感器缺失或组合多变的视觉分割任务。
同时使用来自多个传感器的多模态输入训练分割模型虽直观有益,但实际面临单模态偏见问题——多模态分割器过度依赖特定模态,导致其他模态缺失时性能下降,这在真实应用中常见。为此,我们提出首个可处理任意视觉模态组合的鲁棒分割框架。首先采用并行多模态学习策略训练一个强教师模型;随后,在多尺度表示空间中,通过跨模态和单模态蒸馏,将特征级知识从多模态模型传递至任意模态分割器,以缓解单模态偏见,避免对特定模态的过度依赖。此外,提出一种预测级的模态无关语义蒸馏,实现语义知识的跨模态传递。在合成与真实多传感器基准上的大量实验表明,该方法性能优越。
原文摘要 · Abstract (English)
Simultaneously using multimodal inputs from multiple sensors to train segmentors is intuitively advantageous but practically challenging. A key challenge is unimodal bias, where multimodal segmentors over rely on certain modalities, causing performance drops when others are missing, common in real world applications. To this end, we develop the first framework for learning robust segmentor that can handle any combinations of visual modalities. Specifically, we first introduce a parallel multimodal learning strategy for learning a strong teacher. The cross-modal and unimodal distillation is then achieved in the multi scale representation space by transferring the feature level knowledge from multimodal to anymodal segmentors, aiming at addressing the unimodal bias and avoiding over-reliance on specific modalities. Moreover, a prediction level modality agnostic semantic distillation is proposed to achieve semantic knowledge transferring for segmentation. Extensive experiments on both synthetic and real-world multi-sensor benchmarks demonstrate that our method achieves superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。