arXiv:2412.00857cs.CV2024-12被引 3

用光流引导的高效扩散模型提升视频修复一致性与速度

Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion

  • 双分支结构先恢复光流,再用多尺度光流引导修复
  • 无需训练的隐空间插值加速去噪过程,提升效率
  • 光流注意力缓存机制降低计算开销,适合实时应用

文本引导的视频修复技术显著提升了内容生成性能。近期方法普遍采用扩散模型,虽能实现高质量修复,但在时间一致性与计算效率上仍存在瓶颈。为此,本文提出光流引导的高效扩散框架(FloED),以增强视频连贯性。FloED采用双分支架构:时无关光流分支优先重建受损光流,多尺度光流适配器为修复主分支提供运动指导;同时提出无需训练的隐空间插值方法,利用光流扭曲加速多步去噪;结合光流注意力缓存机制,有效降低光流引入的计算成本。在背景修复与物体移除任务上的大量实验表明,FloED在质量与效率上均优于当前最优扩散方法。代码与模型将公开。

原文摘要 · Abstract (English)

The text-guided video inpainting technique has significantly improved the performance of content generation applications. A recent family for these improvements uses diffusion models, which have become essential for achieving high-quality video inpainting results, yet they still face performance bottlenecks in temporal consistency and computational efficiency. This motivates us to propose a new video inpainting framework using optical Flow-guided Efficient Diffusion (FloED) for higher video coherence. Specifically, FloED employs a dual-branch architecture, where the time-agnostic flow branch restores corrupted flow first, and the multi-scale flow adapters provide motion guidance to the main inpainting branch. Besides, a training-free latent interpolation method is proposed to accelerate the multi-step denoising process using flow warping. With the flow attention cache mechanism, FLoED efficiently reduces the computational cost of incorporating optical flow. Extensive experiments on background restoration and object removal tasks show that FloED outperforms state-of-the-art diffusion-based methods in both quality and efficiency. Our codes and models will be made publicly available.

视频修复扩散模型光流引导高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。