arXiv:2509.08302cs.ROcs.CV2025-09中稿 · IEEE Open Journal …综述被引 14

系统梳理自动驾驶感知基础模型的核心能力与挑战。

Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

  • 按四大核心能力构建新分类体系,聚焦通用知识与时空理解。
  • 指出模型在分布外数据和实时系统中的可靠性瓶颈。
  • 适合关注自动驾驶模型设计与落地的科研与工程人员。

基础模型正重塑自动驾驶感知,推动领域从专用深度学习模型转向可泛化、多任务的通用架构。本综述分析这些模型如何应对泛化性、可扩展性及分布外变化带来的挑战。提出围绕四大核心能力的新分类框架:泛化知识、空间理解、多传感器鲁棒性与时间推理。针对每项能力,系统梳理前沿方法并强调其重要性。不同于传统以方法为中心的综述,本工作以能力为导向,提炼概念设计原则,为模型开发提供清晰指引。最后讨论关键挑战,包括实时系统集成、计算开销以及幻觉与分布外失效等问题,并提出未来研究方向,助力基础模型在自动驾驶中的安全有效部署。

原文摘要 · Abstract (English)

Foundation models are revolutionizing autonomous driving perception, transitioning the field from narrow, task-specific deep learning models to versatile, general-purpose architectures trained on vast, diverse datasets. This survey examines how these models address critical challenges in autonomous perception, including limitations in generalization, scalability, and robustness to distributional shifts. The survey introduces a novel taxonomy structured around four essential capabilities for robust performance in dynamic driving environments: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal reasoning. For each capability, the survey elucidates its significance and comprehensively reviews cutting-edge approaches. Diverging from traditional method-centric surveys, our unique framework prioritizes conceptual design principles, providing a capability-driven guide for model development and clearer insights into foundational aspects. We conclude by discussing key challenges, particularly those associated with the integration of these capabilities into real-time, scalable systems, and broader deployment challenges related to computational demands and ensuring model reliability against issues like hallucinations and out-of-distribution failures. The survey also outlines crucial future research directions to enable the safe and effective deployment of foundation models in autonomous driving systems.

自动驾驶基础模型感知综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。