arXiv:2506.16209cs.CVcs.LG2025-06被引 1

用视频GAN生成车辆轨迹,更真实且推理快。

VideoGAN-based Trajectory Proposal for Automated Vehicles

  • 用鸟瞰图视频训练GAN,生成交通场景动态
  • 100小时训练,单次推理<20毫秒,速度快
  • 轨迹分布贴近真实数据,适合自动驾驶规划

能够生成逼真轨迹是提升道路车辆自动化水平的核心。当前基于规则、模型或传统学习的方法难以有效捕捉未来轨迹的复杂多模态分布。本文研究了在鸟瞰视角(BEV)交通场景视频上训练生成对抗网络(GAN)是否能生成统计准确的轨迹,并正确反映各交通参与者间的空间关系。为此,提出一种利用低分辨率BEV占用栅格视频作为训练数据的流程:通过视频生成模型生成交通场景视频,再结合单帧目标检测与帧间目标匹配提取抽象轨迹数据。特别选择GAN架构以实现相比扩散模型更快的训练与推理速度。在Waymo Open Motion Dataset的真实视频上验证,仅需100 GPU小时训练,推理时间低于20毫秒,生成轨迹在空间与动态参数分布上与真实数据高度一致。

原文摘要 · Abstract (English)

Being able to generate realistic trajectory options is at the core of increasing the degree of automation of road vehicles. While model-driven, rule-based, and classical learning-based methods are widely used to tackle these tasks at present, they can struggle to effectively capture the complex, multimodal distributions of future trajectories. In this paper we investigate whether a generative adversarial network (GAN) trained on videos of bird's-eye view (BEV) traffic scenarios can generate statistically accurate trajectories that correctly capture spatial relationships between the agents. To this end, we propose a pipeline that uses low-resolution BEV occupancy grid videos as training data for a video generative model. From the generated videos of traffic scenarios we extract abstract trajectory data using single-frame object detection and frame-to-frame object matching. We particularly choose a GAN architecture for the fast training and inference times with respect to diffusion models. We obtain our best results within 100 GPU hours of training, with inference times under 20\,ms. We demonstrate the physical realism of the proposed trajectories in terms of distribution alignment of spatial and dynamic parameters with respect to the ground truth videos from the Waymo Open Motion Dataset.

轨迹生成视频生成GAN自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。