arXiv:2409.06189cs.CV2024-09被引 5

让车载相机运动可控,生成多视角驾驶视频更连贯。

MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control

  • 用相机参数作为条件输入,控制视频生成中的镜头运动。
  • 在多视角生成中保持时空一致性,提升视觉连贯性。
  • 适合自动驾驶仿真数据生成,对场景建模有帮助。

高质量的驾驶视频生成对于自动驾驶模型的训练数据至关重要。然而,现有生成模型很少关注多视角任务下车载相机运动的控制能力,而这正是驾驶视频生成的关键。为此,我们提出 MyGo,一个端到端的视频生成框架,将车载相机运动作为条件,以提升相机可控性与多视角一致性。MyGo 通过额外的插件模块将相机参数注入预训练视频扩散模型,尽可能保留预训练模型的丰富知识。同时,在每个视图生成过程中,利用对极约束和邻近视图信息,增强时空一致性。实验结果表明,MyGo 在通用相机控制视频生成及多视角驾驶视频生成任务上均达到当前最优性能,为自动驾驶环境模拟提供了更精准的基础。项目页:https://metadrivescape.github.io/papers_project/MyGo/page.html

原文摘要 · Abstract (English)

High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which is essential for driving video generation. Therefore, we propose MyGo, an end-to-end framework for video generation, introducing motion of onboard cameras as conditions to make progress in camera controllability and multi-view consistency. MyGo employs additional plug-in modules to inject camera parameters into the pre-trained video diffusion model, which retains the extensive knowledge of the pre-trained model as much as possible. Furthermore, we use epipolar constraints and neighbor view information during the generation process of each view to enhance spatial-temporal consistency. Experimental results show that MyGo has achieved state-of-the-art results in both general camera-controlled video generation and multi-view driving video generation tasks, which lays the foundation for more accurate environment simulation in autonomous driving. Project page: https://metadrivescape.github.io/papers_project/MyGo/page.html

视频生成自动驾驶多视角可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。