arXiv:2603.15185cs.ROcs.AI2026-03被引 6

提出轻量级端到端驾驶架构BevAD,提升闭环驾驶的可扩展性与鲁棒性。

What Matters for Scalable and Robust Learning in End-to-End Driving Planners?

  • 采用鸟瞰图特征与轨迹解耦设计,增强感知与规划的协同
  • 在Bench2Drive上实现72.7%成功率,支持纯模仿学习的数据扩展
  • 揭示模块组合对闭环性能的关键影响,适合自动驾驶系统研发者

端到端自动驾驶因在交互场景中学习鲁棒行为并具备数据可扩展性而受到关注。现有架构多基于感知与规划模块分离、通过鸟瞰图特征网格等隐表示连接的方式,保持端到端可微。然而,这类范式主要基于开环数据集,评估不仅关注驾驶表现,还包含中间感知任务。令人遗憾的是,开环表现优异的架构往往难以在闭环驾驶中实现可扩展学习。本文系统重审了三种常见架构模式对闭环性能的影响:(1) 高分辨率感知表示,(2) 轨迹解耦表示,(3) 生成式规划。关键在于,我们评估了这些模式的组合效应,揭示出意外局限与未被充分探索的协同作用。基于此,我们提出BevAD——一种轻量级且高度可扩展的端到端驾驶架构。BevAD在Bench2Drive基准上达到72.7%成功率,并展现出强大的数据扩展能力,仅使用纯模仿学习。代码与模型已公开:https://dmholtz.github.io/bevad/

原文摘要 · Abstract (English)

End-to-end autonomous driving has gained significant attention for its potential to learn robust behavior in interactive scenarios and scale with data. Popular architectures often build on separate modules for perception and planning connected through latent representations, such as bird's eye view feature grids, to maintain end-to-end differentiability. This paradigm emerged mostly on open-loop datasets, with evaluation focusing not only on driving performance, but also intermediate perception tasks. Unfortunately, architectural advances that excel in open-loop often fail to translate to scalable learning of robust closed-loop driving. In this paper, we systematically re-examine the impact of common architectural patterns on closed-loop performance: (1) high-resolution perceptual representations, (2) disentangled trajectory representations, and (3) generative planning. Crucially, our analysis evaluates the combined impact of these patterns, revealing both unexpected limitations as well as underexplored synergies. Building on these insights, we introduce BevAD, a novel lightweight and highly scalable end-to-end driving architecture. BevAD achieves 72.7% success rate on the Bench2Drive benchmark and demonstrates strong data-scaling behavior using pure imitation learning. Our code and models are publicly available here: https://dmholtz.github.io/bevad/

端到端驾驶鸟瞰图可扩展性模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。