让所有视觉模态平等参与分割,动态选择最优组合。
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
- 通过分层模态选择机制,让各模态在不同特征层级自主贡献。
- 在真实与合成数据集上均达顶尖性能,且支持任意模态组合。
- 适合传感器失效或光照复杂场景,对多模态系统设计有启发。
本文针对模态无关语义分割(MaSS)挑战,旨在使每个模态在每个特征粒度上都能发挥价值。在真实世界中,同时训练并有效融合任意组合的视觉模态对鲁棒的多模态融合至关重要,但目前仍研究不足。现有方法常以RGB为中心,将其他模态视为次要,导致结构不对称;而在夜间等场景下,事件数据等模态表现更优。为此,我们提出MAGIC++框架,包含两个可插拔模块:多模态交互模块用于高效处理多模态输入并提取互补信息,采用通道和空间引导;统一多尺度任意模态选择模块则利用聚合特征作为基准,在分层特征空间中基于相似度评分排序各模态特征。该方法在每个特征粒度上摆脱对RGB的依赖,更好应对传感器故障与环境噪声,同时保障分割性能。在常见多模态设置下,MAGIC++在真实与合成基准上均达到最先进水平;在新型模态无关设置中,相比先前方法显著领先。
原文摘要 · Abstract (English)
In this paper, we address the challenging modality-agnostic semantic segmentation (MaSS), aiming at centering the value of every modality at every feature granularity. Training with all available visual modalities and effectively fusing an arbitrary combination of them is essential for robust multi-modal fusion in semantic segmentation, especially in real-world scenarios, yet remains less explored to date. Existing approaches often place RGB at the center, treating other modalities as secondary, resulting in an asymmetric architecture. However, RGB alone can be limiting in scenarios like nighttime, where modalities such as event data excel. Therefore, a resilient fusion model must dynamically adapt to each modality's strengths while compensating for weaker inputs.To this end, we introduce the MAGIC++ framework, which comprises two key plug-and-play modules for effective multi-modal fusion and hierarchical modality selection that can be equipped with various backbone models. Firstly, we introduce a multi-modal interaction module to efficiently process features from the input multi-modal batches and extract complementary scene information with channel-wise and spatial-wise guidance. On top, a unified multi-scale arbitrary-modal selection module is proposed to utilize the aggregated features as the benchmark to rank the multi-modal features based on the similarity scores at hierarchical feature spaces. This way, our method can eliminate the dependence on RGB modality at every feature granularity and better overcome sensor failures and environmental noises while ensuring the segmentation performance. Under the common multi-modal setting, our method achieves state-of-the-art performance on both real-world and synthetic benchmarks. Moreover, our method is superior in the novel modality-agnostic setting, where it outperforms prior arts by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。