平衡视频编辑的连贯性与可编辑性,提升生成质量
Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

- 引入时序Mamba模块,四方向扫描增强帧间一致性
- 基于频域变换的噪声注入策略,提升编辑灵活性
- 在保持输入视频真实性的前提下实现高质量编辑
近期,扩散模型在文本引导视频编辑领域取得显著进展。然而,现有方法通常难以兼顾时间连贯性与可编辑性,二者往往呈负相关。为此,我们提出一种名为EquiEdit的高质量视频编辑框架,协同提升编辑视频的时间连贯性与可编辑性,并实现两者的良好平衡。在时间连贯性方面,设计了具有时序感知能力的Mamba模块,沿四个预设方向扫描融合后的视频序列,有效增强帧间一致性。对于可编辑性,提出基于频域变换的噪声注入策略,利用傅里叶变换保留初始潜在噪声中的隐含结构,确保编辑后视频的帧间一致性和对输入视频的高保真度。大量定性和定量实验表明,该方法在时间连贯性、可编辑性及输入视频保真度方面均表现优异。
原文摘要 · Abstract (English)
Recently, diffusion models have achieved considerable success in the text-guided video editing domain. However, existing works often struggle to balance the trade-off between temporal consistency and editability in video editing, with consistency and editability typically being inversely related. To address this, we propose a high-quality video editing framework enhanced for consistency and editability, named EquiEdit, which improves coordinatively the temporal consistency and editability of the edited videos while achieving a balance between the two. In terms of temporal consistency, the proposed temporal Mamba module with a tailored temporal-aware scanning scans fused video sequences following four designed directions, effectively enhancing the inter-frame consistency of edited videos. For editability, we design a noise injection strategy based on the spectral transformation to increase editing flexibility, where the Fourier transform is used to preserve the hidden structure in the initial latent noise used for editing, ensuring inter-frame consistency of the edited video and fidelity to the input video. Extensive qualitative and quantitative experiments demonstrate the effectiveness of our method in terms of temporal consistency and editability, as well as its great fidelity to the input video itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。