提出新框架让不同模态协同互补,提升跨模态语义分割性能
CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation
- 通过窗口化跨模态交互实现隐式对齐,保留各模态独特优势
- 在多个数据集上显著提升脆弱模态的分割精度,超越基线方法
- 适合需要融合多源感知数据的自动驾驶与机器人场景
模态无关语义分割(MaSS)旨在实现任意输入模态组合下的鲁棒场景理解。现有方法通常依赖显式特征对齐来实现模态同质化,这会削弱各模态的特有优势并破坏其固有互补性。为此,本文提出全新互补学习框架CHARM,通过两个组件实现协同调和而非同质化:(1) 相互感知单元(MPU),通过基于窗口的跨模态交互实现隐式对齐,使各模态互为查询与上下文,发现模态间交互对应关系;(2) 双路径优化策略,将训练解耦为协作学习策略(CoL)用于互补融合,以及个体增强策略(InE)保护模态特异性优化。在多个数据集与骨干网络上的实验表明,CHARM始终优于基线方法,尤其在脆弱模态上取得显著提升。本工作从模型同质化转向协同调和,真正实现多样性的和谐共存。
原文摘要 · Abstract (English)
Modality-agnostic Semantic Segmentation (MaSS) aims to achieve robust scene understanding across arbitrary combinations of input modality. Existing methods typically rely on explicit feature alignment to achieve modal homogenization, which dilutes the distinctive strengths of each modality and destroys their inherent complementarity. To achieve cooperative harmonization rather than homogenization, we propose CHARM, a novel complementary learning framework designed to implicitly align content while preserving modality-specific advantages through two components: (1) Mutual Perception Unit (MPU), enabling implicit alignment through window-based cross-modal interaction, where modalities serve as both queries and contexts for each other to discover modality-interactive correspondences; (2) A dual-path optimization strategy that decouples training into Collaborative Learning Strategy (CoL) for complementary fusion learning and Individual Enhancement Strategy (InE) for protected modality-specific optimization. Experiments across multiple datasets and backbones indicate that CHARM consistently outperform the baselines, with significant increment on the fragile modalities. This work shifts the focus from model homogenization to harmonization, enabling cross-modal complementarity for true harmony in diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。