arXiv:2604.02948cs.CV2026-04

提出跨模态编织框架,实现任意模态组合下的高效语义分割。

CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation

论文配图:CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
图 1 · 摘自论文原文
  • 设计可选的模态交互模块,动态融合多模态信息。
  • 在多个基准上达到顶尖性能,参数增加极少。
  • 适合需要灵活适配新模态组合的场景。

多模态语义分割通过融合不同传感模态的互补信息展现出巨大潜力。然而,现有方法通常依赖精心设计的融合策略,或使用模态特异性适配,或依赖松散耦合的交互,限制了灵活性并导致跨模态协调效果不佳。此外,这些方法在高效信息交换与保持各模态独特性之间难以平衡,尤其在不同模态组合下表现不稳定。为此,我们提出 CrossWeaver,一种适用于任意模态组合的多模态融合框架。其核心是模态交互块(MIB),在编码器中实现选择性且可靠性感知的跨模态交互;轻量级缝合对齐融合(SAF)模块进一步聚合增强特征。在多个多模态语义分割基准上的大量实验表明,该框架以极少额外参数实现了最先进性能,并展现出对未见模态组合的强大泛化能力。

原文摘要 · Abstract (English)

Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use modality-specific adaptations or rely on loosely coupled interactions, thereby limiting flexibility and resulting in less effective cross-modal coordination. Moreover, these methods often struggle to balance efficient information exchange with preserving the unique characteristics of each modality across different modality combinations. To address these challenges, we propose CrossWeaver, a simple yet effective multimodal fusion framework for arbitrary-modality semantic segmentation. Its core is a Modality Interaction Block (MIB), which enables selective and reliability-aware cross-modal interaction within the encoder, while a lightweight Seam-Aligned Fusion (SAF) module further aggregates the enhanced features. Extensive experiments on multiple multimodal semantic segmentation benchmarks demonstrate that our framework achieves state-of-the-art performance with minimal additional parameters and strong generalization to unseen modality combinations.

多模态语义分割融合框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。