端到端自动驾驶正转向以规划为核心的新范式,强调安全与路线合规。
Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms

- 用结构化表示和轨迹输出实现规划导向的端到端驾驶
- 从开环匹配转向闭环仿真与人类偏好评估,提升真实场景适应性
- 适合关注自动驾驶架构演进与评测标准的从业者
端到端自动驾驶已从摄像头到控制的回归模型,演变为以规划为核心的系统,采用结构化表示、轨迹级输出及日益真实的评估协议。本文综述了行为克隆、条件模仿学习、特权蒸馏、BEV与向量规划、统一感知-预测-规划架构、基于世界模型的规划器以及视觉-语言-动作系统的发展。我们指出,现代端到端驾驶的关键区别不在于是否使用中间表示,而在于这些表示是否被学习、监督并用于支持安全、可行且符合路线的规划。为梳理文献,我们沿输入表示、规划输出、监督信号与评估协议四个维度整合现有方法。进一步分析发现,评估范式已从开环轨迹匹配转向闭环仿真、非反应式真实日志评估、长尾测试及人类偏好感知指标。研究显示,缺乏一致基准的架构进步难以解释,仅靠位移为基础的开环指标无法充分证明驾驶的安全性与人类对齐。最后提出不确定性感知规划、学习者-专家差异、运行时安全保障、语言-动作对齐、世界模型验证及可复现基准等开放挑战。
原文摘要 · Abstract (English)
End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imitation learning, privileged distillation, BEV and vectorized planning, unified perception-prediction-planning architectures, world-model-based planners, and vision-language-action systems. We argue that the key distinction in modern end-to-end driving is not whether intermediate representations are used, but whether they are learned, supervised, and evaluated to support safe, feasible, and route-compliant planning. To organize the literature, we synthesize existing methods along four axes: input representation, planning output, supervision signal, and evaluation protocol. We further examine the benchmark shift from open-loop trajectory matching to closed-loop simulation, non-reactive real-log evaluation, long-tail testing, and human-preference-aware metrics. Our analysis highlights that architectural progress is difficult to interpret without benchmark-consistent evaluation, and that displacement-based open-loop metrics alone provide limited evidence for safe and human-aligned driving. We conclude with open challenges in uncertainty-aware planning, learner-expert mismatch, runtime safety assurance, language-action grounding, world-model validation, and reproducible benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。