无需微调或注意力修改,高效实现视频编辑
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
- 通过流变换直接操作潜在空间,避免计算开销
- 推理速度提升20倍,内存减少85%
- 适用于CogVideoX等主流模型,效果领先
视频扩散变换器(Video DiTs)的出现标志着视频生成的重要里程碑。然而,将现有视频编辑方法直接应用于Video DiTs通常带来巨大计算开销,源于资源密集型的注意力修改或微调。为缓解此问题,我们提出DFVEdit,一种专为Video DiTs设计的高效零样本视频编辑方法。DFVEdit通过在干净潜在空间上直接进行流变换,消除对注意力修改和微调的需求。具体而言,我们观察到编辑与采样可在连续流视角下统一。基于此,我们提出条件增量流向量(CDFV)——一种理论无偏的DFV估计,并结合隐式交叉注意力(ICA)引导与嵌入强化(ER)以进一步提升编辑质量。实验表明,相较于基于注意力工程的方法,DFVEdit在Video DiTs上实现至少20倍的推理加速和85%的内存降低。大量定量与定性实验验证,DFVEdit可无缝应用于主流Video DiTs(如CogVideoX和Wan2.1),在结构保真度、时空一致性及编辑质量方面达到当前最优表现。
原文摘要 · Abstract (English)
The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due to resource-intensive attention modification or finetuning. To alleviate this problem, we present DFVEdit, an efficient zero-shot video editing method tailored for Video DiTs. DFVEdit eliminates the need for both attention modification and fine-tuning by directly operating on clean latents via flow transformation. To be more specific, we observe that editing and sampling can be unified under the continuous flow perspective. Building upon this foundation, we propose the Conditional Delta Flow Vector (CDFV) -- a theoretically unbiased estimation of DFV -- and integrate Implicit Cross Attention (ICA) guidance as well as Embedding Reinforcement (ER) to further enhance editing quality. DFVEdit excels in practical efficiency, offering at least 20x inference speed-up and 85% memory reduction on Video DiTs compared to attention-engineering-based editing methods. Extensive quantitative and qualitative experiments demonstrate that DFVEdit can be seamlessly applied to popular Video DiTs (e.g., CogVideoX and Wan2.1), attaining state-of-the-art performance on structural fidelity, spatial-temporal consistency, and editing quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。