arXiv:2505.21205cs.CV2025-05被引 3

提升视频补帧的终点帧约束,避免破坏输入表示

EF-VI: Enhancing End-Frame Injection for Video Inbetweening

  • 用轻量模块仅编码终点帧并生成时序自适应特征注入模型
  • 在多个数据集上显著优于现有基线方法,提升视频连贯性
  • 适合需要高精度终点控制的视频生成任务

视频补帧旨在根据给定的起始帧和终点帧生成中间视频序列。当前最先进方法主要通过直接微调或双向时序采样扩展大规模预训练图像到视频扩散模型(I2V-DMs),但前者导致终点帧约束较弱,后者则不可避免地破坏视频帧的输入表示,影响性能。为在强化终点帧约束的同时避免输入表示被破坏,本文提出专为新型基于Transformer的I2V-DMs设计的视频补帧框架EF-VI。其通过增强注入机制高效加强终点帧约束,核心是提出的轻量级模块EF-Net,该模块仅编码终点帧,并将其扩展为时序自适应的帧级特征,注入I2V-DM中。大量实验表明,与多种基线相比,本方法在多个数据集上表现更优。

原文摘要 · Abstract (English)

Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-trained Image-to-Video Diffusion Models (I2V-DMs) by incorporating the end-frame condition via direct fine-tuning or temporally bidirectional sampling. However, the former results in a weak end-frame constraint, while the latter inevitably disrupts the input representation of video frames, leading to suboptimal performance. To improve the end-frame constraint while avoiding disruption of the input representation, we propose a novel video inbetweening framework specific to recent and more powerful transformer-based I2V-DMs, termed EF-VI. It efficiently strengthens the end-frame constraint by utilizing an enhanced injection. This is based on our proposed well-designed lightweight module, termed EF-Net, which encodes only the end frame and expands it into temporally adaptive frame-wise features injected into the I2V-DM. Extensive experiments demonstrate the superiority of our EF-VI compared with other baselines.

视频生成扩散模型补帧条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。