arXiv:2504.18576cs.RO2025-04被引 22

用轨迹提示生成高保真驾驶场景,提升动态物体时序一致性。

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

  • 将轨迹转为文本提示和2D运动先验,实现精准控制。
  • 在nuScenes和Waymo上超越专用模型,生成视频质量更高。
  • 轻量级运动对齐模块增强长序列动态元素连贯性。

本文提出DriVerse,一种基于单张图像和未来轨迹的生成式驾驶场景模拟模型。以往自动驾驶世界模型直接输入轨迹或离散控制信号,导致控制输入与2D生成模型隐含特征对齐不佳,生成视频保真度低。部分方法使用粗略文本指令或离散车辆控制信号,缺乏精细引导能力,难以评估真实自动驾驶算法。DriVerse引入两种互补的显式轨迹引导:通过预定义趋势词汇表将轨迹编码为文本提示,实现语言无缝融合;将3D轨迹转换为2D空间运动先验,增强对静态场景内容的控制。为更好处理动态物体,进一步设计轻量级运动对齐模块,聚焦动态像素的帧间一致性,显著提升长序列中移动元素的时序连贯性。DriVerse仅需少量训练且无需额外数据,在nuScenes和Waymo数据集上的未来视频生成任务中表现优于专用模型。代码与模型将公开发布。

原文摘要 · Abstract (English)

This paper presents DriVerse, a generative model for simulating navigation-driven driving scenes from a single image and a future trajectory. Previous autonomous driving world models either directly feed the trajectory or discrete control signals into the generation pipeline, leading to poor alignment between the control inputs and the implicit features of the 2D base generative model, which results in low-fidelity video outputs. Some methods use coarse textual commands or discrete vehicle control signals, which lack the precision to guide fine-grained, trajectory-specific video generation, making them unsuitable for evaluating actual autonomous driving algorithms. DriVerse introduces explicit trajectory guidance in two complementary forms: it tokenizes trajectories into textual prompts using a predefined trend vocabulary for seamless language integration, and converts 3D trajectories into 2D spatial motion priors to enhance control over static content within the driving scene. To better handle dynamic objects, we further introduce a lightweight motion alignment module, which focuses on the inter-frame consistency of dynamic pixels, significantly enhancing the temporal coherence of moving elements over long sequences. With minimal training and no need for additional data, DriVerse outperforms specialized models on future video generation tasks across both the nuScenes and Waymo datasets. The code and models will be released to the public.

驾驶模拟轨迹生成运动对齐多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。