无需微调,用噪声结构实现精准视频编辑
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

- 根据编辑区域动态分配噪声水平,优化起始点
- 利用生成模型先验信息提升内容一致性与视觉质量
- 适合追求高效、无训练视频编辑的开发者
视频编辑面临重大挑战。尽管一系列无需微调的方法避免了大量数据收集和模型训练,但往往未能充分利用噪声潜空间中的丰富信息,导致效果不佳。为此,我们提出一种无需微调、基于指令的视频编辑框架。从噪声潜空间视角出发,设计结构化噪声初始化策略(SNIS),通过为编辑区域分配更高噪声水平(促进内容变化),未编辑区域分配更低噪声水平(保持内容一致),确保更优的编辑起点。引入噪声引导机制(NGM),利用生成模型中的视频先验,有效融合噪声潜空间中的丰富信息,指导去噪过程,从而保留未编辑内容并维持整体视觉连贯性。实验表明,所提方法在视觉质量与性能上均优于现有方法,达到当前最优水平。
原文摘要 · Abstract (English)
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit{tuning-free, instruction-based} video editing framework. We approach video editing from the perspective of noisy latent: we design a Structural Noise Initialization Strategy (SNIS) to secure a superior editing starting point by assigning higher noise levels to edited regions (to facilitate content change) and lower noise levels to unedited regions (to maintain content consistency). We introduce a Noise Guidance Mechanism (NGM), which leverages the video prior in the generative model and effectively integrates rich information within the noisy latent to guide the denoising process, thereby preserving unedited content and overall visual coherence. Experiments show that our proposed method achieves better visual quality and state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。