EponaV2通过预测未来3D场景和语义图,实现无需感知的更智能驾驶规划。
EponaV2: Driving World Model with Comprehensive Future Reasoning

- 基于人类驾驶推理思路,预测未来3D几何与语义地图以增强环境理解
- 在三个NAV SIM基准上实现最优性能,轨迹规划评分提升5.5个EPDMS
- 适合研究自动驾驶世界模型、未来预测与端到端规划的学者与工程师
数据规模在追求通用智能中起关键作用。然而,当前自动驾驶普遍采用感知-规划范式,严重依赖昂贵的人工标注来监督轨迹规划,限制了其可扩展性。相比之下,现有无感知驾驶世界模型虽表现优异,但其规划推理仅基于下一帧图像预测,缺乏足够监督,常导致场景理解不全面,轨迹规划效果不佳。本文提出EponaV2,一种新型无感知驾驶世界模型范式,实现高质量且具备综合未来推理能力的规划。受人类驾驶员预判三维结构与语义的启发,模型训练为预测更全面的未来表征,可进一步解码为未来几何与语义地图。提取三维与语义模态使模型深入理解周围环境,未来预测任务显著提升真实世界推理能力,最终优化轨迹规划。此外,借鉴大语言模型训练策略,引入流匹配组相对策略优化机制,进一步提高规划精度。在三个NAV SIM基准上,EponaV2作为无感知模型达到最先进水平(+1.3PDMS,+5.5EPDMS),验证了方法有效性。
原文摘要 · Abstract (English)
Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which severely limits its scalability. Conversely, although existing perception-free driving world models achieve impressive driving performance, their real-world reasoning ability for planning is solely built on next frame image forecasting. Due to the lack of enough supervision, these models often struggle with comprehensive scene understanding, resulting in unsatisfactory trajectory planning. In this paper, we propose EponaV2, a novel paradigm of driving world models, which achieves high-quality planning with comprehensive future reasoning. Inspired by how human drivers anticipate 3D geometry and semantics, we train our model to forecast more comprehensive future representations, which can be additionally decoded to future geometry and semantic maps. Extracting the 3D and semantic modalities enables our model to deeply understand the surrounding environment, and the future prediction task significantly enhances the real-world reasoning capabilities of EponaV2, ultimately leading to improved trajectory planning. Moreover, inspired by the training recipe of Large Language Models (LLMs), we introduce a flow matching group relative policy optimization mechanism to further improve planning accuracy. The state-of-the-art (SOTA) performances of EponaV2 among perception-free models on three NAVSIM benchmarks (+1.3PDMS, +5.5EPDMS) demonstrate the effectiveness of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。