用逆动力学把未来预测转为可执行的驾驶轨迹,提升端到端自动驾驶规划能力
IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving

- 通过逆动力学模型解析未来场景变化中的运动含义,生成规划相关的轨迹增量
- 在NAVSIM v1/v2上实现当前最佳性能,长时序轨迹一致性显著提升
- 适合关注世界模型与轨迹规划融合、追求高精度端到端驾驶系统的研究者
端到端自动驾驶正成为从传感器观测直接学习规划的有效范式,基于世界模型的方法进一步增强了对场景未来演化的显式推理能力。然而,仅预测未来状态并不足以改善规划,除非能将未来演变转化为与规划相关的轨迹更新。现有方法大多仅预测未来场景状态,未显式解码状态转移中隐藏的运动信息,导致未来推理与可执行运动生成之间耦合较弱。为此,本文提出IDOL——一种基于逆动力学引导的未来预测框架,用于潜空间鸟瞰图(latent BEV)下的世界模型端到端规划。IDOL首先通过BEV世界模型预测多个未来潜空间场景状态,再利用逆动力学模型分析相邻未来状态间的转换,解码出与运动相关的轨迹特征,并恢复解释世界演化过程的规划相关运动增量。这些由逆动力学生成的信号被用于优化规划轨迹,使未来预测从被动感知转为主动规划指导。一个轻量级闭环精修模块通过重用优化后的轨迹进行新一轮前瞻推理,进一步提升长时序一致性。实验表明,IDOL在NAVSIM v1和NAVSIM v2基准上优于同类方法,显著增强了世界建模与规划之间的耦合性。
原文摘要 · Abstract (English)
End-to-end autonomous driving has emerged as a compelling paradigm for learning planning directly from sensor observations, while recent world-model-based approaches further enrich this paradigm by enabling explicit reasoning about how the scene may evolve in the future. Yet future prediction alone does not guarantee better planning unless the predicted evolution can be converted into planning-relevant trajectory updates. Many current methods still forecast future scene states without explicitly decoding the motion implications hidden in state transitions. As a result, future reasoning often remains descriptively useful but only weakly coupled to executable motion generation. To address this limitation, we propose \mathbf{IDOL}, an inverse-dynamics-guided future prediction framework for world-model-based end-to-end planning in latent BEV space, where inverse dynamics serves as the key bridge between future prediction and trajectory optimization. IDOL first predicts multiple future latent scene states with a BEV world model, then applies an inverse dynamics model to adjacent latent futures to decode transition-aware trajectory features and recover planning-relevant motion deltas that explain how the latent world evolves over time. These inverse-dynamics-derived signals are used to optimize the planned trajectory, turning future forecasting from passive scene anticipation into actionable planning guidance. A lightweight closed-loop refinement module further improves long-horizon consistency by reusing the optimized trajectory for another round of future-aware reasoning. By introducing inverse dynamics into latent future reasoning, IDOL tightens the coupling between world modeling and planning. Extensive experiments on the NAVSIM v1 and NAVSIM v2 benchmarks show that IDOL achieves state-of-the-art performance among comparable methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。