用新方法让文本生成视频模型更精准地编辑视频内容。
VideoDirector: Precise Video Editing via Text-to-Video Models
- 分离时空信息,提升视频逆向生成精度
- 编辑后视频在动作流畅性与真实感上达顶尖水平
- 适合需要精细局部修改的视频编辑场景
尽管基于文本生成图像(T2I)模型的先逆向再编辑范式已展现良好效果,但直接将其扩展至文本生成视频(T2V)模型仍存在严重伪影,如色彩闪烁和内容失真。因此,现有视频编辑方法仍主要依赖T2I模型,而其缺乏时序一致性生成能力,常导致编辑质量下降。本文将该问题归因于:1)空间-时间紧密耦合,原始关键帧逆向策略难以解耦视频扩散模型中的时空信息;2)复杂的时空布局,原始交叉注意力控制无法有效保留未编辑内容。为此,我们提出空间-时间解耦引导(STDG)与多帧空文本优化策略,提供更精确的关键帧时序提示,实现更精准的逆向生成;同时引入自注意力控制机制,提升局部内容编辑的保真度。实验表明,所提方法(VideoDirector)充分挖掘了T2V模型强大的时序生成能力,在准确性、运动平滑性、真实感及未编辑内容保真度方面均达到当前最优水平。
原文摘要 · Abstract (English)
Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods primarily rely on T2I models, which inherently lack temporal-coherence generative ability, often resulting in inferior editing results. In this paper, we attribute the failure of the typical editing paradigm to: 1) Tightly Spatial-temporal Coupling. The vanilla pivotal-based inversion strategy struggles to disentangle spatial-temporal information in the video diffusion model; 2) Complicated Spatial-temporal Layout. The vanilla cross-attention control is deficient in preserving the unedited content. To address these limitations, we propose a spatial-temporal decoupled guidance (STDG) and multi-frame null-text optimization strategy to provide pivotal temporal cues for more precise pivotal inversion. Furthermore, we introduce a self-attention control strategy to maintain higher fidelity for precise partial content editing. Experimental results demonstrate that our method (termed VideoDirector) effectively harnesses the powerful temporal generation capabilities of T2V models, producing edited videos with state-of-the-art performance in accuracy, motion smoothness, realism, and fidelity to unedited content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。