无需训练,用少量步骤实现高质量视频编辑。
StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation

- 从噪声生成数据视角重构视频编辑,保留快速采样能力。
- 在少步采样下超越现有方法,时间成本极低且效果稳定。
- 适合追求高效可控视频编辑的研究者与开发者。
现有视频编辑方法通常需多次迭代,仍难以保证高质量结果。我们归因于普遍采用的数据到数据范式,与现代生成模型不兼容。为此,本文从噪声到数据视角出发,提出无需训练的流式生成视频编辑方法 StreamEdit,保留少步采样优势的同时无缝注入源视频条件。基于预训练流式生成模型,StreamEdit 引入双分支快速采样、自注意力桥接与交叉注意力对齐/增强机制,满足采样与条件约束。进一步提出面向源视频的引导策略提升目标生成质量,并设计视觉提示策略增强编辑灵活性与实用性。大量实验表明,StreamEdit 在多种视频编辑任务中持续优于现有方法,即使在少步设置下也表现优异,且耗时极低。代码与结果详见:https://dsl-lab.github.io/StreamEdit/
原文摘要 · Abstract (English)
Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet satisfying editing results. We attribute this limitation to the prevalent data-to-data paradigm, which is less compatible with modern generative models than noise-to-data generation. To address this gap, we revisit video editing from a noise-to-data perspective and propose Streaming-Generation-based Video Editing (StreamEdit), which preserves few-step sampling while seamlessly injecting source-video conditions. Built on pre-trained streaming generation models, StreamEdit introduces dual-branch fast sampling with a self-attention bridge and cross-attention grounding/boosting to satisfy both sampling and conditioning requirements. We further propose source-oriented guidance to improve target-generation quality, and a visual prompting strategy to enhance editing flexibility and practicality. The method is effective, robust, and generalizable across different models. Extensive experiments on diverse video editing tasks show that StreamEdit consistently outperforms existing approaches, even in few-step settings with minimal time cost. Code and results are available at: https://dsl-lab.github.io/StreamEdit/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。