arXiv:2502.08244cs.CV2025-02CVPR被引 41

用光流实现可控制的自然视频生成,无需真实相机参数

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

  • 用光流同时表征相机与物体运动,实现精准控制
  • 无需真实相机参数即可训练,支持任意视频输入
  • 利用背景光流实现细节级相机动作控制,适合影视生成

我们提出FloVD,一种用于相机可控视频生成的新视频扩散模型。该模型利用光流表示相机和运动物体的运动。这一方法带来两大优势:由于光流可直接从视频中估计,本方法可使用任意训练视频而无需真实相机参数;此外,背景光流编码了不同视角间的三维相关性,使方法能通过背景运动实现精细相机控制。为在支持精细相机控制的同时合成自然物体运动,框架采用两阶段视频生成流程:光流生成与光流条件下的视频合成。大量实验表明,该方法在相机控制精度与自然物体运动生成方面优于先前方法。

原文摘要 · Abstract (English)

We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can be directly estimated from videos, our approach allows for the use of arbitrary training videos without ground-truth camera parameters. Moreover, as background optical flow encodes 3D correlation across different viewpoints, our method enables detailed camera control by leveraging the background motion. To synthesize natural object motion while supporting detailed camera control, our framework adopts a two-stage video synthesis pipeline consisting of optical flow generation and flow-conditioned video synthesis. Extensive experiments demonstrate the superiority of our method over previous approaches in terms of accurate camera control and natural object motion synthesis.

视频生成扩散模型光流相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。