无需反演,直接在数据空间引导视频生成,实现精准文本编辑。
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
- 不依赖反演,用微分方程直接引导视频沿时空流形演化。
- 在多个基准上实现最佳指令遵循与运动一致性表现。
- 适合需要高保真、无训练视频编辑的开发者与研究者。
文本驱动视频编辑旨在根据自然语言指令修改视频内容。尽管近期无训练方法利用预训练扩散模型,但通常依赖反演-编辑范式,即先将视频映射到隐空间再编辑。然而,反演过程不完全准确,常导致外观保真度下降和运动一致性受损。为此,我们提出FlowDirector,一种全新的无训练、无反演视频编辑框架。该框架将编辑过程建模为数据空间中的直接演化,通过常微分方程(ODE)引导视频沿其固有时空流形平滑过渡,从而避免不准确的反演步骤。在此基础上,我们引入三种流校正策略:1)方向感知流校正增强与源方向相反的分量并移除无关项,打破保守流线,实现更强的结构与纹理变化;2)运动-外观解耦在每个时间步将运动一致性作为能量项优化,显著提升一致性与运动迁移效果;3)差异平均引导策略利用多个候选流的差异,在低成本下近似低方差状态,抑制伪影并稳定轨迹。在多种编辑任务与基准上的大量实验表明,FlowDirector在指令遵循、时序一致性和背景保留方面均达到最先进水平,建立了无需反演的连贯视频编辑新范式。
原文摘要 · Abstract (English)
Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm maps the video to a latent space before editing. However, the inversion process is not perfectly accurate, often compromising appearance fidelity and motion consistency. To address this, we introduce FlowDirector, a novel training-free and inversion-free video editing framework. Our framework models the editing process as a direct evolution in the data space. It guides the video to transition smoothly along its inherent spatio-temporal manifold using an ordinary differential equation (ODE), thereby avoiding the inaccurate inversion step. From this foundation, we introduce three flow correction strategies for appearance, motion, and stability: 1) Direction-aware flow correction amplifies components that oppose the source direction and removes irrelevant terms, breaking conservative streamlines and enabling stronger structural and textural changes. 2) Motion-appearance decoupling optimizes motion agreement as an energy term at each timestep, significantly improving consistency and motion transfer. 3) Differential averaging guidance strategy leverages differences among multiple candidate flows to approximate a low variance regime at low cost, suppressing artifacts and stabilizing the trajectory. Extensive experiments across various editing tasks and benchmarks demonstrate that FlowDirector achieves state-of-the-art performance in instruction following, temporal consistency, and background preservation, establishing an efficient new paradigm for coherent video editing without inversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。