arXiv:2603.08240cs.CV2026-03中稿 · ICLR

让多模态感知在单模态失效时仍能稳定工作

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

  • 通过自适应融合机制,动态处理缺失模态的特征
  • 在激光雷达失效时仍保持性能,避免感知崩溃
  • 适合自动驾驶等对可靠性要求高的场景

协同感知通过整合多智能体视角提升感知范围并克服遮挡问题。现有多模态方法依赖互补传感器提升性能,但一旦关键传感器(如激光雷达)失效,系统极易失败。根本原因在于特征融合导致单模态特征与下游模块间语义不匹配。本文首次在协同感知领域提出单模态可操作的多模态协同感知(SiMO),采用长度自适应多模态融合(LAMMA)机制,在模态故障时自适应处理剩余模态特征,同时保持语义空间一致性。此外,提出创新的‘预训练-对齐-融合-重建’训练策略,缓解模态竞争问题,确保各模态分支独立性。实验表明,SiMO有效对齐多模态特征,同时保留模态特异性特征,可在所有单模态下维持最优性能。

原文摘要 · Abstract (English)

Collaborative perception integrates multi-agent perspectives to enhance the sensing range and overcome occlusion issues. While existing multimodal approaches leverage complementary sensors to improve performance, they are highly prone to failure--especially when a key sensor like LiDAR is unavailable. The root cause is that feature fusion leads to semantic mismatches between single-modality features and the downstream modules. This paper addresses this challenge for the first time in the field of collaborative perception, introducing Single-Modality-Operable Multimodal Collaborative Perception (SiMO). By adopting the proposed Length-Adaptive Multi-Modal Fusion (LAMMA), SiMO can adaptively handle remaining modal features during modal failures while maintaining consistency of the semantic space. Additionally, leveraging the innovative "Pretrain-Align-Fuse-RD" training strategy, SiMO addresses the issue of modality competition--generally overlooked by existing methods--ensuring the independence of each individual modality branch. Experiments demonstrate that SiMO effectively aligns multimodal features while simultaneously preserving modality-specific features, enabling it to maintain optimal performance across all individual modalities. The implementation details can be found in https://github.com/dempsey-wen/SiMO.

多模态感知自动驾驶鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。