arXiv:2509.09828cs.CVcs.LG2025-09被引 6

用深度信息动态调整多传感器融合,提升自动驾驶语义感知鲁棒性

DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception

  • 基于深度引导的跨模态融合,通过局部深度令牌动态调节融合策略
  • 在MUSES和DeLiVER数据集上达到最新最优的全景与语义分割性能
  • 适合需要高鲁棒性感知的自动驾驶系统研发人员参考

自动驾驶中的鲁棒语义感知依赖于有效融合具有互补优势与缺陷的多传感器数据。现有先进融合方法对输入空间各处的数据处理方式一致,难以应对复杂环境。为此,我们提出一种新型深度引导的多模态融合方法,通过引入激光雷达提供的深度信息实现条件感知融合的升级。DGFusion将多模态分割建模为多任务问题,利用激光雷达数据作为模型输入及深度监督信号。其辅助深度头学习深度感知特征,生成空间可变的局部深度令牌,结合全局条件令牌动态调节跨模态注意力机制。该设计使融合策略能根据场景中各传感器的可靠性(主要由深度决定)自适应调整。此外,我们设计了一种针对稀疏且噪声大的激光雷达输入的鲁棒损失函数。实验表明,该方法在挑战性的MUSES和DeLiVER数据集上均取得当前最佳的全景与语义分割结果。代码与模型已开源。

原文摘要 · Abstract (English)

Robust semantic perception for autonomous vehicles relies on effectively combining multiple sensors with complementary strengths and weaknesses. State-of-the-art sensor fusion approaches to semantic perception often treat sensor data uniformly across the spatial extent of the input, which hinders performance when faced with challenging conditions. By contrast, we propose a novel depth-guided multimodal fusion method that upgrades condition-aware fusion by integrating depth information. Our network, DGFusion, poses multimodal segmentation as a multi-task problem, utilizing the lidar measurements, which are typically available in outdoor sensor suites, both as one of the model's inputs and as ground truth for learning depth. Our corresponding auxiliary depth head helps to learn depth-aware features, which are encoded into spatially varying local depth tokens that condition our attentive cross-modal fusion. Together with a global condition token, these local depth tokens dynamically adapt sensor fusion to the spatially varying reliability of each sensor across the scene, which largely depends on depth. In addition, we propose a robust loss for our depth, which is essential for learning from lidar inputs that are typically sparse and noisy in adverse conditions. Our method achieves state-of-the-art panoptic and semantic segmentation performance on the challenging MUSES and DeLiVER datasets. Code and models are available at https://github.com/timbroed/DGFusion

传感器融合自动驾驶深度引导语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。