新评估指标能更好预测真实自动驾驶表现,但仍有局限。
Do Open-Loop Metrics Predict Closed-Loop Driving? A Cross-Benchmark Correlation Study of NAVSIM and Bench2Drive
- 用最新安全导向的开环指标(NAVSIM)跨基准对比
- 进展度(EP)是闭环表现最强预测因子,碰撞率不是
- 简单三指标公式即可媲美复杂指标,适合快速筛选模型
开环评估可快速复现自动驾驶规划器性能,但其能否预测真实闭环驾驶表现仍存疑。以往研究表明,传统开环指标如平均位移误差(ADE)和最终位移误差(FDE)与闭环驾驶得分无可靠相关性。本文系统比对了15种先进方法在NAVSIM(开环)和Bench2Drive(闭环)上的结果,构建了8组完整配对数据。分析发现:(1) NAVSIM PDM总分与闭环驾驶得分呈强正相关但非单调,存在排名反转;(2) 单个子指标中,自车进展度(EP)是最佳预测项,显著优于安全关键的碰撞指标NC;(3) 开闭环中安全与进展权衡机制不同:过度追求安全导致进展不足的方法在NAVSIM中排名高,但在闭环中因超时和低速被惩罚。进一步表明,仅用3个指标即可达到与完整5指标相同的预测力(斯皮尔曼ρ=0.90,n=8),说明当前先进方法中,到达时间(TTC)与舒适度已趋饱和,新增信息有限。此外,小偏差累积成大失败的雪球效应可能是残余差距的原因。
原文摘要 · Abstract (English)
Open-loop evaluation offers fast, reproducible assessment of autonomous driving planners, but its ability to predict real closed-loop driving performance remains questionable. Prior work has shown that traditional open-loop metrics such as Average Displacement Error (ADE) and Final Displacement Error (FDE) exhibit no reliable correlation with closed-loop Driving Score. In this paper, we ask whether the more recent, safety-aware open-loop metrics introduced by NAVSIM~v2 can bridge this gap. By systematically cross-referencing published results from 15 state-of-the-art methods across NAVSIM (open-loop) and Bench2Drive (closed-loop), we compile a paired dataset of open-loop sub-metrics and closed-loop performance, yielding 8 methods with complete paired data. Our analysis reveals three key findings: (1) the aggregate NAVSIM PDM Score shows a strong positive but non-monotonic correlation with Bench2Drive Driving Score, with clear ranking inversions; (2) among individual NAVSIM sub-metrics, Ego Progress (EP) is the strongest single predictor of closed-loop success, substantially exceeding the safety-critical collision metric NC; (3) the safety-progress trade-off manifests differently in open-loop and closed-loop: methods that maximize safety at the expense of progress rank highly in NAVSIM but underperform in closed-loop due to timeout and slow-driving penalties. We further demonstrate that a much simpler 3-metric formula matches the predictive power of the full 5-metric PDMS at the same Spearman $ρ{=}0.90$ on our paired sample of $n{=}8$ methods, suggesting that within current state-of-the-art methods -- where TTC and Comfort approach saturation -- these two sub-metrics add little marginal information for closed-loop ranking. Additionally, we identify the snowball effect -- where small open-loop deviations compound into closed-loop failures -- as a candidate mechanism for the residual gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。