arXiv:2503.09151cs.CVcs.AI2025-03ICCV被引 38

用视频转视频方式生成多视角同步视频,无需大规模4D数据训练

Reangle-A-Video: 4D Video Generation as Video-to-Video Translation

  • 将多视角视频生成视为视频到视频的翻译任务,利用现有图像和视频扩散模型
  • 通过自监督微调学习跨视角不变运动,实现多视角动作同步
  • 适合需要动态相机控制或视点迁移的视频生成场景

我们提出Reangle-A-Video,一种从单个输入视频生成同步多视角视频的统一框架。不同于主流方法在大规模4D数据上训练多视角视频扩散模型,我们的方法将多视角视频生成任务重新定义为视频到视频的翻译,利用公开可用的图像和视频扩散先验。该方法分为两个阶段:(1) 多视角运动学习:通过自监督方式同步微调图像到视频扩散变换器,从一组形变视频中提炼出视角无关的运动信息;(2) 多视角一致图像到图像翻译:在推理时使用DUSt3R进行跨视角一致性引导,将输入视频的第一帧形变并修复为多个相机视角下的起始图像。在静态视角传输和动态相机控制任务上的大量实验表明,Reangle-A-Video优于现有方法,为多视角视频生成提供了新方案。代码与数据将公开发布。

原文摘要 · Abstract (English)

We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the multi-view video generation task as video-to-videos translation, leveraging publicly available image and video diffusion priors. In essence, Reangle-A-Video operates in two stages. (1) Multi-View Motion Learning: An image-to-video diffusion transformer is synchronously fine-tuned in a self-supervised manner to distill view-invariant motion from a set of warped videos. (2) Multi-View Consistent Image-to-Images Translation: The first frame of the input video is warped and inpainted into various camera perspectives under an inference-time cross-view consistency guidance using DUSt3R, generating multi-view consistent starting images. Extensive experiments on static view transport and dynamic camera control show that Reangle-A-Video surpasses existing methods, establishing a new solution for multi-view video generation. We will publicly release our code and data. Project page: https://hyeonho99.github.io/reangle-a-video/

视频生成多视角扩散模型视频翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。