无需反演即可高效适配流模型,实现创意视频编辑。
Taming Flow-based I2V Models for Creative Video Editing
- 不依赖反演,通过向量场修正融合源视频信息。
- 在多个数据集上优于现有方法,保持视频一致性。
- 轻量级插件式设计,适合快速部署于新模型。
尽管图像编辑技术已取得显著进展,但根据用户意图操控视频的视频编辑仍面临挑战。现有基于图像条件的视频编辑方法通常需要特定模型设计的反演或大量优化,限制了其利用最新基于流匹配的图像到视频(I2V)模型将图像编辑能力迁移至视频领域的潜力。为此,我们提出 IF-V2V,一种无需反演的方法,可无需显著计算开销地适配现成的流匹配型 I2V 模型进行视频编辑。为规避反演,我们设计了带样本偏差的向量场修正,通过在去噪向量场中引入偏差项,将源视频信息融入去噪过程。为进一步以模型无关方式确保与源视频的一致性,我们提出结构与运动保持初始化,生成包含结构信息的运动感知时序相关噪声。此外,我们还引入偏差缓存机制,在几乎不增加计算成本的前提下最小化去噪向量修正的开销,且不影响编辑质量。评估表明,该方法在编辑质量和一致性方面均优于现有方法,提供了一种轻量级、即插即用的视觉创意实现方案。
原文摘要 · Abstract (English)
Although image editing techniques have advanced significantly, video editing, which aims to manipulate videos according to user intent, remains an emerging challenge. Most existing image-conditioned video editing methods either require inversion with model-specific design or need extensive optimization, limiting their capability of leveraging up-to-date image-to-video (I2V) models to transfer the editing capability of image editing models to the video domain. To this end, we propose IF-V2V, an Inversion-Free method that can adapt off-the-shelf flow-matching-based I2V models for video editing without significant computational overhead. To circumvent inversion, we devise Vector Field Rectification with Sample Deviation to incorporate information from the source video into the denoising process by introducing a deviation term into the denoising vector field. To further ensure consistency with the source video in a model-agnostic way, we introduce Structure-and-Motion-Preserving Initialization to generate motion-aware temporally correlated noise with structural information embedded. We also present a Deviation Caching mechanism to minimize the additional computational cost for denoising vector rectification without significantly impacting editing quality. Evaluations demonstrate that our method achieves superior editing quality and consistency over existing approaches, offering a lightweight plug-and-play solution to realize visual creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。