综述自动驾驶轨迹规划中的基础模型进展与挑战
Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- 构建统一分类框架,系统梳理37种轨迹规划基础模型
- 揭示多模态模型在感知-决策端到端能力上的优势
- 关注开源情况,助力研究者快速复现与选型
多模态基础模型的兴起显著改变了自动驾驶技术范式,推动从传统手工设计转向统一的基础模型方法,能够直接从原始传感器输入推断运动轨迹。这类方法还可融合自然语言作为额外模态,以视觉-语言-动作(VLA)模型为代表。本文通过统一分类体系,全面评估相关方法的架构设计、方法优势及其内在能力与局限性。涵盖37种近期提出的基于基础模型的轨迹规划方法,并评估其源代码与数据集开放程度,为研究者和从业者提供实用参考。配套网页已上线,按分类整理所有方法:https://github.com/fiveai/FMs-for-driving-trajectories
原文摘要 · Abstract (English)
The emergence of multi-modal foundation models has markedly transformed the technology for autonomous driving, shifting away from conventional and mostly hand-crafted design choices towards unified, foundation-model-based approaches, capable of directly inferring motion trajectories from raw sensory inputs. This new class of methods can also incorporate natural language as an additional modality, with Vision-Language-Action (VLA) models serving as a representative example. In this review, we provide a comprehensive examination of such methods through a unifying taxonomy to critically evaluate their architectural design choices, methodological strengths, and their inherent capabilities and limitations. Our survey covers 37 recently proposed approaches that span the landscape of trajectory planning with foundation models. Furthermore, we assess these approaches with respect to the openness of their source code and datasets, offering valuable information to practitioners and researchers. We provide an accompanying webpage that catalogues the methods based on our taxonomy, available at: https://github.com/fiveai/FMs-for-driving-trajectories
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。