系统梳理多传感器融合感知技术,覆盖从基础方法到大模型融合的全链条进展。
A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects
- 按任务无关视角整合多模态、多视角、时序及大模型融合方法
- 涵盖3D目标检测、语义分割等下游任务的融合技术体系
- 适合研究具身智能、自动驾驶等领域的学者快速掌握技术全景
多传感器融合感知(MSFP)是具身人工智能的关键技术,可支持3D目标检测、语义分割等下游任务及自动驾驶、集群机器人等应用场景。现有综述多聚焦单一任务或领域,且仅从多模态融合角度出发,缺乏对多视图融合、时序融合等多样方法的全面覆盖。本文从任务无关视角出发,系统梳理MSFP背景,回顾多模态与多智能体融合方法,分析时序融合技术,并探讨大语言模型时代下的多模态融合新范式。最后讨论开放挑战与未来方向。本综述旨在帮助研究人员理解MSFP的重要进展,为后续研究提供参考。
原文摘要 · Abstract (English)
Multi-sensor fusion perception (MSFP) is a key technology for embodied AI, which can serve a variety of downstream tasks (e.g., 3D object detection and semantic segmentation) and application scenarios (e.g., autonomous driving and swarm robotics). Recently, impressive achievements on AI-based MSFP methods have been reviewed in relevant surveys. However, we observe that the existing surveys have some limitations after a rigorous and detailed investigation. For one thing, most surveys are oriented to a single task or research field, such as 3D object detection or autonomous driving. Therefore, researchers in other related tasks often find it difficult to benefit directly. For another, most surveys only introduce MSFP from a single perspective of multi-modal fusion, while lacking consideration of the diversity of MSFP methods, such as multi-view fusion and time-series fusion. To this end, in this paper, we hope to organize MSFP research from a task-agnostic perspective, where methods are reported from various technical views. Specifically, we first introduce the background of MSFP. Next, we review multi-modal and multi-agent fusion methods. A step further, time-series fusion methods are analyzed. In the era of LLM, we also investigate multimodal LLM fusion methods. Finally, we discuss open challenges and future directions for MSFP. We hope this survey can help researchers understand the important progress in MSFP and provide possible insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。