arXiv:2506.01061cs.CV2025-06中稿 · IEEE Transactions …综述被引 14

系统梳理250+篇视频插帧论文,解析主流方法与未来方向

AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation

  • 按核心设计原则分类,区分中心时间与任意时间插帧范式
  • 涵盖从传统运动补偿到扩散模型的全链条方法演进
  • 适合想快速掌握该领域全景的研究者与工程师

视频帧插值(VFI)是核心低层视觉任务,旨在合成现有帧之间的中间帧,同时保证空间和时间一致性。过去几十年中,VFI方法已从基于运动补偿的传统手段发展为涵盖核、光流、混合、相位、GAN、Transformer、Mamba及最新扩散模型等多样化的深度学习方法。本文提出AceVFI,对超过250篇代表性论文进行系统性综述。基于核心设计原理与架构特征对方法进行分类,并将其划分为两大学习范式:中心时间帧插值(CTFI)与任意时间帧插值(ATFI)。分析了大运动、遮挡、光照变化和非线性运动等关键挑战。回顾标准数据集、损失函数与评估指标,并探讨其在其他领域的应用,展望未来研究方向。本综述旨在为研究人员与实践者提供现代VFI领域的全面参考。

原文摘要 · Abstract (English)

Video Frame Interpolation (VFI) is a core low-level vision task that synthesizes intermediate frames between existing ones while ensuring spatial and temporal coherence. Over the past decades, VFI methodologies have evolved from classical motion compensation-based approach to a wide spectrum of deep learning-based approaches, including kernel-, flow-, hybrid-, phase-, GAN-, Transformer-, Mamba-, and most recently, diffusion-based models. We introduce AceVFI, a comprehensive and up-to-date review of the VFI field, covering over 250 representative papers. We systematically categorize VFI methods based on their core design principles and architectural characteristics. Further, we classify them into two major learning paradigms: Center-Time Frame Interpolation (CTFI) and Arbitrary-Time Frame Interpolation (ATFI). We analyze key challenges in VFI, including large motion, occlusion, lighting variation, and non-linear motion. In addition, we review standard datasets, loss functions, evaluation metrics. We also explore VFI applications in other domains and highlight future research directions. This survey aims to serve as a valuable reference for researchers and practitioners seeking a thorough understanding of the modern VFI landscape.

视频插帧深度学习综述计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。