arXiv:2601.11475cs.CV2026-01被引 5

用语言引导生成未来交通场景,让自动驾驶模型更懂意图、更稳推理。

Generative Scenario Rollouts for End-to-End Autonomous Driving

  • 通过自回归滚动生成未来场景和动作响应,实现多智能体协同规划。
  • 在Bench2Drive上驾驶得分提升15.7%,成功率提高26.2%,表现领先。
  • 支持零样本鲁棒性,适合需要可解释与长时序决策的自动驾驶系统。

视觉-语言-动作(VLA)模型正成为端到端自动驾驶中高效的规划模型。然而,现有方法多依赖稀疏轨迹标注进行模仿学习,未能充分发挥其生成能力。本文提出生成场景滚动(GeRo),一种即插即用的VLA框架,通过自回归滚动策略联合执行规划与语言对齐的未来交通场景生成。首先,模型在规划、运动与语言任务监督下,将本车与其它智能体动态编码为隐状态令牌,实现文本对齐生成。随后,GeRo基于多视角图像、场景描述及本车动作提问,生成未来隐状态令牌与文本回复,指导长时程滚动。引入滚动一致性损失,利用真实或伪标签稳定预测,缓解漂移并保持文本-动作对齐。该设计使GeRo实现时间一致、语言驱动的滚动,支持长时程推理与多智能体规划。在Bench2Drive上,GeRo驾驶得分提升15.7%,成功率提升26.2%。结合强化学习,实现开环与闭环最优性能,展现强零样本鲁棒性。结果表明,生成式语言引导推理是构建更安全、可解释自动驾驶系统的有效基础。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize their potential as generative models. We propose Generative Scenario Rollouts (GeRo), a plug-and-play framework for VLA models that jointly performs planning and generation of language-grounded future traffic scenes through an autoregressive rollout strategy. First, a VLA model is trained to encode ego vehicle and agent dynamics into latent tokens under supervision from planning, motion, and language tasks, facilitating text-aligned generation. Next, GeRo performs language-conditioned autoregressive generation. Given multi-view images, a scenario description, and ego-action questions, it generates future latent tokens and textual responses to guide long-horizon rollouts. A rollout-consistency loss stabilizes predictions using ground truth or pseudo-labels, mitigating drift and preserving text-action alignment. This design enables GeRo to perform temporally consistent, language-grounded rollouts that support long-horizon reasoning and multi-agent planning. On Bench2Drive, GeRo improves driving score and success rate by +15.7 and +26.2, respectively. By integrating reinforcement learning with generative rollouts, GeRo achieves state-of-the-art closed-loop and open-loop performance, demonstrating strong zero-shot robustness. These results highlight the promise of generative, language-conditioned reasoning as a foundation for safer and more interpretable end-to-end autonomous driving.

自动驾驶生成模型多智能体语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。