arXiv:2503.15875cs.CV2025-03被引 18

MiLA生成长达1分钟的高保真自动驾驶视频,解决动态场景长期生成误差问题。

MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving

  • 采用粗到精分阶段生成策略,稳定视频输出并纠正动态物体失真。
  • 在nuScenes数据集上生成视频质量达当前最优,支持最长60秒连续画面。
  • 适合自动驾驶系统训练数据增强,尤其适用于复杂动态场景模拟。

近年来,数据驱动技术推动了自动驾驶系统发展,但稀有且多样化的训练数据仍需大量设备与人力投入。世界模型通过预测和生成未来环境状态,可合成带标注的视频数据用于训练,是潜在解决方案。然而,现有方法在动态场景中难以生成长时间、一致的视频,易累积误差。为此,我们提出MiLA框架,能够生成长达一分钟的高保真视频。MiLA采用粗到精(Coarse-to-Re(fine))策略,在稳定生成过程的同时修正动态物体形变。此外,引入时间渐进去噪调度器与联合去噪校正流模块,进一步提升生成质量。在nuScenes数据集上的大量实验表明,MiLA在视频生成质量上达到当前最优水平。更多信息请访问项目官网:https://github.com/xiaomi-mlab/mila.github.io。

原文摘要 · Abstract (English)

In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant investment in equipment and labor. World models, which predict and generate future environmental states, offer a promising solution by synthesizing annotated video data for training. However, existing methods struggle to generate long, consistent videos without accumulating errors, especially in dynamic scenes. To address this, we propose MiLA, a novel framework for generating high-fidelity, long-duration videos up to one minute. MiLA utilizes a Coarse-to-Re(fine) approach to both stabilize video generation and correct distortion of dynamic objects. Additionally, we introduce a Temporal Progressive Denoising Scheduler and Joint Denoising and Correcting Flow modules to improve the quality of generated videos. Extensive experiments on the nuScenes dataset show that MiLA achieves state-of-the-art performance in video generation quality. For more information, visit the project website: https://github.com/xiaomi-mlab/mila.github.io.

自动驾驶视频生成世界模型长时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。