arXiv:2508.07560cs.ROcs.CV2025-08综述被引 4

系统梳理自动驾驶中鸟瞰图感知的三阶段演进与安全挑战

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

  • 从单车单模态到多车协同,分三阶段构建安全感知框架
  • 覆盖多种数据集,评估其在复杂场景下的可靠性表现
  • 聚焦开放世界难题,适合研究自动驾驶安全的学者参考

鸟瞰图(BEV)感知已成为自动驾驶的基础范式,支持统一的空间表征,促进多传感器融合与多智能体协作。随着自动驾驶从受控环境向真实道路部署过渡,如何在遮挡、恶劣天气和动态交通等复杂场景下保障BEV感知的安全性与可靠性仍是关键挑战。本文首次从安全关键视角全面综述了BEV感知技术,系统分析了三个渐进阶段:单模态车载感知、多模态车载感知以及多智能体协同感知。此外,我们考察了涵盖车载、路侧及协同设置的公开数据集,评估其对安全性和鲁棒性的适配度。文中识别出若干开放世界挑战,包括开集识别、大规模未标注数据、传感器退化及智能体间通信延迟,并提出未来研究方向,如与端到端自动驾驶系统融合、具身智能以及大语言模型的应用。

原文摘要 · Abstract (English)

Bird's-Eye-View (BEV) perception has become a foundational paradigm in autonomous driving, enabling unified spatial representations that support robust multi-sensor fusion and multi-agent collaboration. As autonomous vehicles transition from controlled environments to real-world deployment, ensuring the safety and reliability of BEV perception in complex scenarios - such as occlusions, adverse weather, and dynamic traffic - remains a critical challenge. This survey provides the first comprehensive review of BEV perception from a safety-critical perspective, systematically analyzing state-of-the-art frameworks and implementation strategies across three progressive stages: single-modality vehicle-side, multimodal vehicle-side, and multi-agent collaborative perception. Furthermore, we examine public datasets encompassing vehicle-side, roadside, and collaborative settings, evaluating their relevance to safety and robustness. We also identify key open-world challenges - including open-set recognition, large-scale unlabeled data, sensor degradation, and inter-agent communication latency - and outline future research directions, such as integration with end-to-end autonomous driving systems, embodied intelligence, and large language models.

自动驾驶鸟瞰图安全感知多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。