arXiv:2503.19776cs.CV2025-03CVPR被引 28

多模态专家融合让自动驾驶传感器在故障下仍稳定感知

Resilient Sensor Fusion under Adverse Sensor Failures via Multi-Modal Expert Fusion

  • 用三个独立专家解码器分离相机与激光雷达依赖
  • 自适应路由选择最优专家,提升极端故障下的检测精度
  • 适合自动驾驶在恶劣环境或传感器失效时的鲁棒感知需求

现代自动驾驶感知系统依赖激光雷达(LiDAR)与摄像头等多模态传感器。尽管融合架构提升了复杂环境性能,但在严重传感器故障(如激光雷达束减少、丢失、视野受限、摄像头失效或遮挡)下仍表现显著下降,根源在于现有框架中的模态间依赖。本文提出高效鲁棒的LiDAR-摄像头3D目标检测模型MoME,采用多专家融合策略。其通过三个并行专家解码器分别处理仅相机特征、仅激光雷达特征或两者结合的特征,实现模态完全解耦。引入多专家解码(MED)框架,每个目标查询由自适应查询路由(AQR)动态选择最合适的专家解码器,依据相机与激光雷达特征质量进行决策。该机制确保每条查询由最优专家处理,从而在多样传感器故障场景中保持稳健性能。在nuScenes-R基准上评估,MoME在极端天气和传感器故障条件下均达到当前最佳表现,显著优于现有模型。

原文摘要 · Abstract (English)

Modern autonomous driving perception systems utilize complementary multi-modal sensors, such as LiDAR and cameras. Although sensor fusion architectures enhance performance in challenging environments, they still suffer significant performance drops under severe sensor failures, such as LiDAR beam reduction, LiDAR drop, limited field of view, camera drop, and occlusion. This limitation stems from inter-modality dependencies in current sensor fusion frameworks. In this study, we introduce an efficient and robust LiDAR-camera 3D object detector, referred to as MoME, which can achieve robust performance through a mixture of experts approach. Our MoME fully decouples modality dependencies using three parallel expert decoders, which use camera features, LiDAR features, or a combination of both to decode object queries, respectively. We propose Multi-Expert Decoding (MED) framework, where each query is decoded selectively using one of three expert decoders. MoME utilizes an Adaptive Query Router (AQR) to select the most appropriate expert decoder for each query based on the quality of camera and LiDAR features. This ensures that each query is processed by the best-suited expert, resulting in robust performance across diverse sensor failure scenarios. We evaluated the performance of MoME on the nuScenes-R benchmark. Our MoME achieved state-of-the-art performance in extreme weather and sensor failure conditions, significantly outperforming the existing models across various sensor failure scenarios.

传感器融合自动驾驶鲁棒检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。