通过像素运动建模,实现复杂镜头轨迹下视频生成的稳定一致。
MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- 将相机与物体运动统一为像素级运动流进行学习。
- 在多个镜头轨迹任务中显著优于当前最优方法。
- 适合需要精准镜头控制的视频生成场景。
基于镜头轨迹生成视频面临一致性与泛化性的挑战,尤其当相机与物体同时运动时。现有方法常分别学习两类运动,易导致相对运动混淆。为此,本文提出新方法,将相机与物体运动统一转换为对应像素的运动。利用稳定的扩散网络,有效学习与指定镜头轨迹相关的参考运动图。结合提取的语义物体先验,输入图像到视频网络,生成能准确跟随指定镜头轨迹且保持物体运动一致的视频。大量实验表明,该模型显著优于当前最优方法。
原文摘要 · Abstract (English)
Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt to learn these motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera and object motions by converting them into the motion of corresponding pixels. Utilizing a stable diffusion network, we effectively learn reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining consistent object motions. Extensive experiments verify that our model outperforms SOTA methods by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。