arXiv:2510.26125cs.CVcs.AI2025-10被引 75

聚焦罕见复杂场景,构建端到端自动驾驶新基准与评估方法。

WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

  • 专为罕见长尾场景设计,含4021段高精度驾驶片段。
  • 提出评分反馈指标RFS,更贴近人类对行车轨迹的偏好判断。
  • 适合研究鲁棒性、泛化能力的自动驾驶系统开发者。

基于视觉的端到端(E2E)自动驾驶在研究界备受关注,因其具备可扩展性并能与多模态大语言模型协同。然而,现有E2E驾驶基准大多仅涵盖常规场景,难以充分检验系统真实潜力。此外,传统开环评估指标常无法有效捕捉驾驶行为的多模态特性或在长尾场景中准确评估性能。为此,我们推出面向端到端自动驾驶的Waymo开放数据集(WOD-E2E),包含4,021段驾驶片段(约12小时),均来自日常生活中发生频率低于0.03%的挑战性长尾场景。每段数据包含高阶路径信息、车辆状态及8个环绕摄像头的360°图像。为评估此类长尾场景下的E2E驾驶表现,我们提出新型开环评估指标:评分反馈得分(RFS)。该指标不比较预测与实际轨迹距离,而是衡量预测轨迹与人工标注的偏好轨迹之间的匹配度。我们已公开全部验证集的评分偏好标签,测试集标签用于2025年WOD-E2E挑战赛。本工作旨在推动具备泛化性、鲁棒性和安全性的端到端自动驾驶智能体的研究发展。

原文摘要 · Abstract (English)

Vision-based end-to-end (E2E) driving has garnered significant interest in the research community due to its scalability and synergy with multimodal large language models (MLLMs). However, current E2E driving benchmarks primarily feature nominal scenarios, failing to adequately test the true potential of these systems. Furthermore, existing open-loop evaluation metrics often fall short in capturing the multi-modal nature of driving or effectively evaluating performance in long-tail scenarios. To address these gaps, we introduce the Waymo Open Dataset for End-to-End Driving (WOD-E2E). WOD-E2E contains 4,021 driving segments (approximately 12 hours), specifically curated for challenging long-tail scenarios that that are rare in daily life with an occurring frequency of less than 0.03%. Concretely, each segment in WOD-E2E includes the high-level routing information, ego states, and 360-degree camera views from 8 surrounding cameras. To evaluate the E2E driving performance on these long-tail situations, we propose a novel open-loop evaluation metric: Rater Feedback Score (RFS). Unlike conventional metrics that measure the distance between predicted way points and the logs, RFS measures how closely the predicted trajectory matches rater-annotated trajectory preference labels. We have released rater preference labels for all WOD-E2E validation set segments, while the held out test set labels have been used for the 2025 WOD-E2E Challenge. Through our work, we aim to foster state of the art research into generalizable, robust, and safe end-to-end autonomous driving agents capable of handling complex real-world situations.

自动驾驶长尾场景评估指标端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。