用分层扩散模型高效生成视频插帧的光流,速度超快精度更高。
Hierarchical Flow Diffusion for Efficient Frame Interpolation
- 分层扩散建模双向光流,缩小去噪搜索空间。
- 插帧精度达当前最优,速度比同类方法快10倍以上。
- 适合追求高速高质视频生成的研究者与开发者。
当前基于扩散的方法在视频帧插值任务中,准确率和效率仍显著落后于非扩散方法。多数方法直接在高维隐空间进行去噪,因隐空间过大而效果不佳。本文提出通过分层扩散模型显式建模双目光流,使去噪过程搜索空间大幅缩小。基于该光流扩散模型,进一步设计光流引导的图像合成器生成最终结果,并实现端到端联合训练。实验表明,本方法在准确率上达到最新水平,且推理速度比其他扩散方法快10倍以上。
原文摘要 · Abstract (English)
Most recent diffusion-based methods still show a large gap compared to non-diffusion methods for video frame interpolation, in both accuracy and efficiency. Most of them formulate the problem as a denoising procedure in latent space directly, which is less effective caused by the large latent space. We propose to model bilateral optical flow explicitly by hierarchical diffusion models, which has much smaller search space in the denoising procedure. Based on the flow diffusion model, we then use a flow-guided images synthesizer to produce the final result. We train the flow diffusion model and the image synthesizer end to end. Our method achieves state of the art in accuracy, and 10+ times faster than other diffusion-based methods. The project page is at: https://hfd-interpolation.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。