无需反演的视频编辑,让多物体运动更稳定
FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing

- 用空间注意力和自适应强度调节稳定视频编辑信号
- 在多物体、快动作场景中保持时序一致性和编辑精度
- 适合需要高效稳定视频生成的研究者与开发者
我们提出 FlowAnchor,一个无需训练的框架,实现稳定高效的无反演流式视频编辑。现有无反演图像编辑方法能高效保留结构并直接调控采样轨迹,但扩展至视频仍面临挑战,尤其在多物体场景或帧数增多时易失效。根源在于高维视频潜在空间中编辑信号的不稳定性,源于空间定位不准及长度引起的强度衰减。FlowAnchor 显式锚定编辑位置与强度,引入空间感知注意力精炼,确保文本引导与空间区域一致对齐;同时采用自适应幅度调制,动态保持足够的编辑强度。两者协同稳定编辑信号,引导流式演化至目标分布。大量实验表明,FlowAnchor 在复杂多物体与快速运动场景下实现了更忠实、时序连贯且计算高效的视频编辑。
原文摘要 · Abstract (English)
We propose FlowAnchor, a training-free framework for stable and efficient inversion-free, flow-based video editing. Inversion-free editing methods have recently shown impressive efficiency and structure preservation in images by directly steering the sampling trajectory with an editing signal. However, extending this paradigm to videos remains challenging, often failing in multi-object scenes or with increased frame counts. We identify the root cause as the instability of the editing signal in high-dimensional video latent spaces, which arises from imprecise spatial localization and length-induced magnitude attenuation. To overcome this challenge, FlowAnchor explicitly anchors both where to edit and how strongly to edit. It introduces Spatial-aware Attention Refinement, which enforces consistent alignment between textual guidance and spatial regions, and Adaptive Magnitude Modulation, which adaptively preserves sufficient editing strength. Together, these mechanisms stabilize the editing signal and guide the flow-based evolution toward the desired target distribution. Extensive experiments demonstrate that FlowAnchor achieves more faithful, temporally coherent, and computationally efficient video editing across challenging multi-object and fast-motion scenarios. The project page is available at https://cuc-mipg.github.io/FlowAnchor.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。