arXiv:2506.21885cs.CVcs.MM2025-06中稿 · IEEE IV 2025综述被引 22

综述智能驾驶多模态传感器融合技术,助力复杂环境感知

Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles

  • 按数据、特征、决策三层分类,系统梳理深度学习融合方法
  • 分析主流多模态数据集,评估其在恶劣天气与城市场景中的表现
  • 探讨视觉-语言模型与端到端自动驾驶中的融合新趋势

多传感器融合在自动驾驶中至关重要,可克服单一传感器局限,实现对环境的全面感知。本文首先将多传感器融合策略形式化为数据级、特征级和决策级三类,并系统回顾了对应各层级的深度学习方法。文中介绍了关键的多模态数据集,讨论其在应对恶劣天气条件和复杂城市环境等现实挑战中的适用性。此外,还探索了视觉-语言模型(VLMs)、大语言模型(LLMs)的融合应用,以及传感器融合在端到端自动驾驶中的作用,强调其提升系统适应性与鲁棒性的潜力。本工作为当前多传感器融合方法及未来方向提供了重要洞察。

原文摘要 · Abstract (English)

Multi-sensor fusion plays a critical role in enhancing perception for autonomous driving, overcoming individual sensor limitations, and enabling comprehensive environmental understanding. This paper first formalizes multi-sensor fusion strategies into data-level, feature-level, and decision-level categories and then provides a systematic review of deep learning-based methods corresponding to each strategy. We present key multi-modal datasets and discuss their applicability in addressing real-world challenges, particularly in adverse weather conditions and complex urban environments. Additionally, we explore emerging trends, including the integration of Vision-Language Models (VLMs), Large Language Models (LLMs), and the role of sensor fusion in end-to-end autonomous driving, highlighting its potential to enhance system adaptability and robustness. Our work offers valuable insights into current methods and future directions for multi-sensor fusion in autonomous driving.

传感器融合自动驾驶多模态VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。