提出新离线评估指标,显著提升预测模型性能与实际驾驶表现的相关性。
Scalable Offline Metrics for Autonomous Driving
- 基于认知不确定性设计新型离线评估指标,捕捉闭环场景中潜在错误
- 相比旧指标,相关性提升超13%,在真实场景中效果更优
- 适用于自动驾驶策略评估,尤其适合关注安全性的研究者
基于感知的规划模型在机器人系统(如自动驾驶汽车)中的真实世界评估可通过离线方式进行,即在预采集的带真值标注的验证数据集上计算模型预测误差。然而,将离线性能外推到在线场景仍面临挑战:看似微小的误差可能累积导致测试时违规或碰撞。这一关系尚未被充分研究,尤其在多样化的闭环评估指标和复杂城市驾驶操作中。本文通过大规模实验重新审视策略评估中的这一被低估的问题。基于仿真分析,我们发现离线与在线表现之间的相关性比以往研究报道的更为薄弱,质疑了现有评估方法的有效性。随后,我们提出一种基于认知不确定性的离线指标,旨在捕捉闭环设置中可能导致错误的事件。该指标相较之前方法相关性提升超过13%。进一步在真实场景中验证,其泛化能力更强,收益更大。
原文摘要 · Abstract (English)
Real-world evaluation of perception-based planning models for robotic systems, such as autonomous vehicles, can be safely and inexpensively conducted offline, i.e. by computing model prediction error over a pre-collected validation dataset with ground-truth annotations. However, extrapolating from offline model performance to online settings remains a challenge. In these settings, seemingly minor errors can compound and result in test-time infractions or collisions. This relationship is understudied, particularly across diverse closed-loop metrics and complex urban maneuvers. In this work, we revisit this undervalued question in policy evaluation through an extensive set of experiments across diverse conditions and metrics. Based on analysis in simulation, we find an even worse correlation between offline and online settings than reported by prior studies, casting doubts on the validity of current evaluation practices and metrics for driving policies. Next, we bridge the gap between offline and online evaluation. We investigate an offline metric based on epistemic uncertainty, which aims to capture events that are likely to cause errors in closed-loop settings. The resulting metric achieves over 13% improvement in correlation compared to previous offline metrics. We further validate the generalization of our findings beyond the simulation environment in real-world settings, where even greater gains are observed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。