arXiv:2506.07713cs.CVcs.AI2025-06被引 7

用光流驱动图像生成视频,实现更连贯的视频编辑。

Consistent Video Editing as Flow-Driven Image-to-Video Generation

  • 将视频编辑拆解为首帧编辑与条件图像到视频生成
  • 通过模拟匹配形变的伪光流序列,提升时序一致性
  • 在多对象和人像编辑上表现优异,适合复杂运动场景

随着视频扩散模型的发展,视频编辑等下游应用得以高效实现,计算成本低。但该任务的核心挑战在于从源视频到编辑后视频的运动迁移过程,需兼顾形状形变建模与时间序列一致性。现有方法难以处理复杂运动模式,且主要局限于物体替换,对非刚性运动如多对象及人像编辑关注不足。本文发现光流在复杂运动建模中具有潜力,提出FlowV2V,将视频编辑重新定义为光流驱动的图像到视频(I2V)生成任务。具体地,该方法将流程分解为首帧编辑与条件I2V生成,并模拟与形变一致的伪光流序列,从而保障编辑过程中的时序一致性。在DAVIS-EDIT数据集上的实验显示,相比现有最先进方法,FlowV2V在DOVER指标上提升13.67%,在形变误差上降低50.66%,显著改善了生成视频的时间一致性和样本质量。此外,通过全面消融实验,分析了首帧范式与光流对齐机制在方法内部的作用。

原文摘要 · Abstract (English)

With the prosper of video diffusion models, down-stream applications like video editing have been significantly promoted without consuming much computational cost. One particular challenge in this task lies at the motion transfer process from the source video to the edited one, where it requires the consideration of the shape deformation in between, meanwhile maintaining the temporal consistency in the generated video sequence. However, existing methods fail to model complicated motion patterns for video editing, and are fundamentally limited to object replacement, where tasks with non-rigid object motions like multi-object and portrait editing are largely neglected. In this paper, we observe that optical flows offer a promising alternative in complex motion modeling, and present FlowV2V to re-investigate video editing as a task of flow-driven Image-to-Video (I2V) generation. Specifically, FlowV2V decomposes the entire pipeline into first-frame editing and conditional I2V generation, and simulates pseudo flow sequence that aligns with the deformed shape, thus ensuring the consistency during editing. Experimental results on DAVIS-EDIT with improvements of 13.67% and 50.66% on DOVER and warping error illustrate the superior temporal consistency and sample quality of FlowV2V compared to existing state-of-the-art ones. Furthermore, we conduct comprehensive ablation studies to analyze the internal functionalities of the first-frame paradigm and flow alignment in the proposed method.

视频编辑光流扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。