arXiv:2507.05173cs.CV2025-07被引 2

基于首尾帧与文本提示生成任意长度的视频中间帧,提升可控性与一致性。

Semantic Frame Interpolation

  • 提出新任务SFI,支持多帧率、带文本控制的语义插帧。
  • 构建SemFi模型,用混合LoRA模块确保不同帧长下内容一致性。
  • 推出首个SFI专用数据集SFI-300K及多维度评估体系,适合视频生成研究者。

基于给定首尾帧与文本提示生成任意长度的中间视频内容,具有重要研究与应用价值。然而,传统帧插值任务多限于少量帧、无文本控制且首尾帧差异小。近期基于Wan的大模型虽具备帧间生成能力,但仅支持固定帧数,对特定帧长效果不佳,且缺乏明确定义与基准。本文首次从学术定义出发,提出新的通用语义帧插值(Semantic Frame Interpolation, SFI)任务,涵盖前述两种场景并支持多帧率推理。为此,我们基于Wan2.1提出SemFi模型,引入混合LoRA模块,确保在不同帧长限制下生成内容高度一致且符合控制条件。同时,我们构建了首个专用于SFI的任务数据集SFI-300K,从SFI视角收集与处理数据,设计多维度评估指标,涵盖图像与视频质量、一致性与多样性等。在SFI-300K上的大量实验表明,所提方法能有效满足SFI任务需求。

原文摘要 · Abstract (English)

Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significant research and application potential. However, traditional frame interpolation tasks primarily focus on scenarios with a small number of frames, no text control, and minimal differences between the first and last frames. Recent community developers have utilized large video models represented by Wan to endow frame-to-frame capabilities. However, these models can only generate a fixed number of frames and often fail to produce satisfactory results for certain frame lengths, while this setting lacks a clear official definition and a well-established benchmark. In this paper, we first propose a new practical Semantic Frame Interpolation (SFI) task from the perspective of academic definition, which covers the above two settings and supports inference at multiple frame rates. To achieve this goal, we propose a novel SemFi model building upon Wan2.1, which incorporates a Mixture-of-LoRA module to ensure the generation of high-consistency content that aligns with control conditions across various frame length limitations. Furthermore, we propose SFI-300K, the first general-purpose dataset and benchmark specifically designed for SFI. To support this, we collect and process data from the perspective of SFI, carefully designing evaluation metrics and methods to assess the model's performance across multiple dimensions, encompassing image and video, and various aspects, including consistency and diversity. Through extensive experiments on SFI-300K, we demonstrate that our method is particularly well-suited to meet the requirements of the SFI task.

视频生成语义插帧大模型多帧率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。