arXiv:2607.07844cs.ROcs.AI2026-07中稿 · 2026 IEEE/RSJ Inte…

提出新基准,测试自动驾驶规划器在陌生城市和执行扰动下的泛化与鲁棒性。

Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning

论文配图:Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning
图 1 · 摘自论文原文
  • 构建双轨测试框架,分别评估语义分布与状态分布的迁移能力。
  • 模仿学习在陌生城市中表现差,尤其在行人密集区;强化学习更耐扰动。
  • 适用于评估规划器真实部署潜力,尤其关注安全与持续进展能力。

尽管基于大规模对象级数据集(如nuPlan)训练的闭环运动规划器在分布内(ID)表现优异,但其在新城市拓扑结构下的泛化能力及执行扰动后的恢复机制仍缺乏研究。为此,我们提出Shift & Drift基准,从两个关键维度严格检验规划器:(1)语义迁移轨道通过新转换流程将DeepScenario Open 3D的航拍数据转为nuPlan仿真环境,实现零样本评估——将北美与新加坡训练的规划器应用于覆盖德国四城及美国旧金山共1,182个场景,包含密集行人-骑行者交互;(2)状态分布漂移轨道注入随机扰动,量化规划器对累积执行误差的鲁棒性。系统评估多种规划范式在语义与状态分布偏移下的失效模式。尽管模仿学习在ID基准中得分高,但在语义迁移下显著失败,尤其在行人密集区,且面对时序相关执行噪声时持续漂移;而强化学习规划器表现出更平滑退化,在两轨道均保持更高安全性和进展指标。研究揭示了模仿保真度与闭环韧性间的权衡关系,为社区提供可靠部署评估基准。

原文摘要 · Abstract (English)

While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, their generalization to novel urban topologies and recovery mechanisms following execution perturbations remain under-explored. To address this, we present Shift & Drift, a novel dual-track benchmark designed to rigorously stress-test motion planners across two critical axes of distribution shift: (1) The Semantic Shift Track leverages a novel conversion pipeline that transforms the aerial, DeepScenario Open 3D dataset into the nuPlan simulation framework. This enables zero-shot evaluation of planners trained on North American and Singaporean data against 1,182 scenarios spanning four German cities and the US city of San Francisco featuring dense pedestrian-cyclist interactions. (2) The State-Distribution Drift Track injects stochastic perturbations into the ego vehicle's dynamics to quantify robustness against compounding execution errors. Based on this, we systematically evaluate the failure modes of diverse planning paradigms under semantic and state-distribution shifts. While imitation learning methods achieve high scores in ID benchmarks, they exhibit significant failures under semantic shift, particularly in pedestrian-dense environments, and suffer from persistent drift when subjected to temporally correlated actuation noise. In contrast, the evaluated reinforcement-learning-based planner demonstrates more graceful degradation, maintaining higher safety and progress metrics across both tracks. Our findings reveal an empirical trade-off between imitation fidelity and closed-loop resilience, providing the community with a rigorous benchmark to evaluate progress toward reliable deployment.

自动驾驶运动规划泛化性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。