arXiv:2512.18241cs.CV2025-12

用语义引导提升流式插帧质量,实现实时高画质视频生成

SG-RIFE: Semantic-Guided Real-Time Intermediate Flow Estimation with Diffusion-Competitive Perceptual Quality

  • 在预训练RIFE模型上注入语义先验,实现参数高效微调
  • 在SNU-FILM上FID/LPIPS优于扩散模型LDMVFI,接近Consec. BB效果
  • 适合追求实时性与高质量视频生成的工业应用

实时视频帧插值长期由基于光流的方法主导,如RIFE,虽具备高吞吐量,但在大运动和遮挡等复杂场景下表现不佳。而近期基于扩散模型的方法(如Consec. BB)虽达到顶尖感知质量,但延迟过高,难以用于实时场景。为此,我们提出语义引导的RIFE(SG-RIFE)。不从头训练,而是通过参数高效的微调策略,将冻结的DINOv3视觉变换器中的语义先验注入预训练的RIFE主干网络。设计了分频敏投影模块(Split-FAPM)压缩并优化高维特征,以及可变形语义融合模块(DSF)对齐语义先验与像素级运动场。在SNU-FILM上的实验表明,语义注入显著提升感知保真度。SG-RIFE在FID/LPIPS指标上优于扩散模型LDMVFI,且在复杂基准上质量接近Consec. BB,同时运行速度显著更快,证明语义一致性使流式方法能在近实时条件下达到扩散模型级别的感知质量。

原文摘要 · Abstract (English)

Real-time Video Frame Interpolation (VFI) has long been dominated by flow-based methods like RIFE, which offer high throughput but often fail in complicated scenarios involving large motion and occlusion. Conversely, recent diffusion-based approaches (e.g., Consec. BB) achieve state-of-the-art perceptual quality but suffer from prohibitive latency, rendering them impractical for real-time applications. To bridge this gap, we propose Semantic-Guided RIFE (SG-RIFE). Instead of training from scratch, we introduce a parameter-efficient fine-tuning strategy that augments a pre-trained RIFE backbone with semantic priors from a frozen DINOv3 Vision Transformer. We propose a Split-Fidelity Aware Projection Module (Split-FAPM) to compress and refine high-dimensional features, and a Deformable Semantic Fusion (DSF) module to align these semantic priors with pixel-level motion fields. Experiments on SNU-FILM demonstrate that semantic injection provides a decisive boost in perceptual fidelity. SG-RIFE outperforms diffusion-based LDMVFI in FID/LPIPS and achieves quality comparable to Consec. BB on complex benchmarks while running significantly faster, proving that semantic consistency enables flow-based methods to achieve diffusion-competitive perceptual quality in near real-time.

视频插帧扩散模型实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。