让自动驾驶提前想象未来场景,决策更前瞻。
See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

- 用预训练世界模型生成未来画面,再基于想象做决策。
- 在NAVSIM和nuScenes上表现优于现有最先进方法。
- 适合研究前瞻式自动驾驶与智能决策系统的学者。
当前端到端自动驾驶规划器本质上是反应式的:它们基于历史和当前观测预测未来动作。我们提出,自动驾驶智能体应先想象未来场景再行动,如同人类驾驶员在行动前会心理模拟“接下来会发生什么”。为此,我们引入了以基础世界模型为核心的规划框架ForeSight,将自动驾驶重新定义为前瞻式决策。不同于将世界模型视为辅助组件,ForeSight将未来场景想象作为动作预测的核心驱动力。该方法分两阶段进行:(1) 通过预训练世界模型生成合理的未来视觉场景;(2) 基于这些想象中的未来进行动作规划。这一从“我现在该做什么?”到“会发生什么,我该如何应对?”的范式转变,实现了真正意义上的前瞻而非反应式规划。通过基于预期情境而非仅当前观测做决策,ForeSight在动态交互场景中表现更优。在NAVSIM和nuScenes上的大量实验表明,显式未来想象显著优于以往最先进方法,验证了其前瞻性驱动的有效性。
原文摘要 · Abstract (English)
Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before deciding, just as human drivers mentally simulate ``what will happen next" before acting. We introduce ForeSight, a foundation world model centric planning framework that reframes autonomous driving as anticipatory decision-making. Rather than treating world models as auxiliary components, ForeSight makes future scene imagination the primary driver of action prediction. Our approach operates in two stages: (1) generating plausible future visual worlds via a pretrained world model, and (2) planning actions conditioned on these imagined futures. This paradigm shift from ``what should I do now?" to ``what will happen, and how should I respond?" enables genuinely anticipatory rather than reactive planning. By grounding decisions in anticipated contexts rather than present observations alone, ForeSight navigates dynamic, interactive scenarios more effectively. Extensive experiments on NAVSIM and nuScenes demonstrate that explicit future imagination significantly outperforms previous state-of-the-art alternatives, validating our foresight-driven approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。