arXiv:2412.17042cs.CV2024-12

用图像生成视频模型做大运动插帧,效果更优。

Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation

  • 设计条件编码器适配图像转视频模型用于插帧
  • 双分支特征提取+跨帧注意力提升大运动处理能力
  • 在FVD指标上优于现有方法,适合大运动场景

近年来视频生成模型发展迅速,本文采用大规模图像到视频扩散模型进行视频帧插值。提出一种条件编码器,将图像到视频模型适配用于大运动帧插值。为提升性能,引入双分支特征提取器,并设计跨帧注意力机制,有效捕捉时空信息,实现中间帧的精准生成。在弗雷歇视频距离(Fréchet Video Distance, FVD)指标上,该方法相比其他先进方法表现更优,尤其在大运动场景下展现出生成式方法的新进展。

原文摘要 · Abstract (English)

With the development of video generation models has advanced significantly in recent years, we adopt large-scale image-to-video diffusion models for video frame interpolation. We present a conditional encoder designed to adapt an image-to-video model for large-motion frame interpolation. To enhance performance, we integrate a dual-branch feature extractor and propose a cross-frame attention mechanism that effectively captures both spatial and temporal information, enabling accurate interpolations of intermediate frames. Our approach demonstrates superior performance on the Fréchet Video Distance (FVD) metric when evaluated against other state-of-the-art approaches, particularly in handling large motion scenarios, highlighting advancements in generative-based methodologies.

视频插帧扩散模型大运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。