arXiv:2503.09215cs.CVcs.AI2025-03AAAI被引 10

统一自车与他车轨迹,实现更真实的驾驶场景视频生成。

Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space

  • 将自车与他车轨迹投影到图像坐标,实现精准匹配控制。
  • 在视频潜在空间中统一建模,生成符合轨迹的动态驾驶画面。
  • 可预测未见场景,适合自动驾驶仿真与策略评估使用。

先进的端到端自动驾驶系统需预测他车运动并规划自车轨迹。现有的世界模型多聚焦自车轨迹,难以控制他车行为,限制了驾驶交互的真实模拟。本文提出EOT-WM驱动世界模型,将自车与他车轨迹统一于视频潜在空间以实现驾驶仿真。针对贝叶斯鸟瞰图(BEV)空间中多轨迹与视频中车辆匹配难题,我们通过像素位置将轨迹投影至图像坐标进行对齐。随后利用时空变分自编码器编码轨迹视频,使其在空间与时间上与驾驶视频潜在表示一致。进一步设计轨迹注入式扩散Transformer,基于轨迹引导去噪潜在表示以生成视频。此外,提出基于控制潜在相似性的评估指标。在nuScenes数据集上的大量实验表明,该模型在FID上优于现有方法30%,在FVD上提升55%。模型还能生成自车产生新轨迹后的未见驾驶场景。

原文摘要 · Abstract (English)

Advanced end-to-end autonomous driving systems predict other vehicles' motions and plan ego vehicle's trajectory. The world model that can foresee the outcome of the trajectory has been used to evaluate the autonomous driving system. However, existing world models predominantly emphasize the trajectory of the ego vehicle and leave other vehicles uncontrollable. This limitation hinders their ability to realistically simulate the interaction between the ego vehicle and the driving scenario. In this paper, we propose a driving World Model named EOT-WM, unifying Ego-Other vehicle Trajectories in videos for driving simulation. Specifically, it remains a challenge to match multiple trajectories in the BEV space with each vehicle in the video to control the video generation. We first project ego-other vehicle trajectories in the BEV space into the image coordinate for vehicle-trajectory match via pixel positions. Then, trajectory videos are encoded by the Spatial-Temporal Variational Auto Encoder to align with driving video latents spatially and temporally in the unified visual space. A trajectory-injected diffusion Transformer is further designed to denoise the noisy video latents for video generation with the guidance of ego-other vehicle trajectories. In addition, we propose a metric based on control latent similarity to evaluate the controllability of trajectories. Extensive experiments are conducted on the nuScenes dataset, and the proposed model outperforms the state-of-the-art method by 30% in FID and 55% in FVD. The model can also predict unseen driving scenes with self-produced trajectories.

自动驾驶世界模型轨迹生成视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。