EDEN通过增强扩散模型提升大运动视频插帧质量
EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
- 用Transformer tokenizer生成精细化中间帧潜空间表示
- 引入时序注意力与起止帧差异嵌入,提升动态运动生成
- 在DAVIS/SNU-FILM上LPIPS降低近10%,适合大运动场景
处理复杂或非线性运动模式一直是视频插帧的难题。尽管基于扩散的方法相比传统光流方法有所改进,但在大运动场景下仍难以生成清晰且时间一致的帧。为此,我们提出EDEN(Enhanced Diffusion for high-quality large-motion video frame interpolation)。该方法首先使用基于Transformer的tokenizer生成扩散模型所需的精细中间帧潜空间表示;随后,在扩散Transformer中引入跨过程的时序注意力机制,并结合起止帧差异嵌入以引导动态运动生成。大量实验表明,EDEN在多个主流基准上达到领先性能:在DAVIS和SNU-FILM上LPIPS降低近10%,在DAIN-HD上提升8%。
原文摘要 · Abstract (English)
Handling complex or nonlinear motion patterns has long posed challenges for video frame interpolation. Although recent advances in diffusion-based methods offer improvements over traditional optical flow-based approaches, they still struggle to generate sharp, temporally consistent frames in scenarios with large motion. To address this limitation, we introduce EDEN, an Enhanced Diffusion for high-quality large-motion vidEo frame iNterpolation. Our approach first utilizes a transformer-based tokenizer to produce refined latent representations of the intermediate frames for diffusion models. We then enhance the diffusion transformer with temporal attention across the process and incorporate a start-end frame difference embedding to guide the generation of dynamic motion. Extensive experiments demonstrate that EDEN achieves state-of-the-art results across popular benchmarks, including nearly a 10% LPIPS reduction on DAVIS and SNU-FILM, and an 8% improvement on DAIN-HD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。