arXiv:2506.01454cs.CV2025-06被引 1

无需训练即可生成高帧率流畅视频,解决动态场景画面闪烁问题。

DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion

  • 利用低帧率视频关键帧,通过滑动窗口去噪实现高帧率生成
  • 在快速运动场景中保持时间连贯性,避免画面闪烁和质量下降
  • 无需微调预训练模型,适用于虚拟现实与游戏等实时应用

扩散模型的进展已显著推动视频生成,能产出高质量且时序一致的视频。然而,在快速运动场景中,生成高帧率(FPS)视频仍面临画面闪烁与长序列质量退化等挑战。现有方法普遍存在计算效率低下、长期视频质量难以保障的问题。本文提出一种全新的免训练高帧率视频生成方法 DiffuseSlide,其核心是利用低帧率视频的关键帧,结合噪声重注入与滑动窗口潜在去噪技术,在不需额外微调的前提下实现平滑、一致的视频输出。大量实验表明,该方法显著提升视频质量,增强时间连贯性与空间保真度,兼具计算高效性与任务适应性,适用于虚拟现实、游戏及高质量内容创作等场景。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have revolutionized video generation, enabling the creation of high-quality, temporally consistent videos. However, generating high frame-rate (FPS) videos remains a significant challenge due to issues such as flickering and degradation in long sequences, particularly in fast-motion scenarios. Existing methods often suffer from computational inefficiencies and limitations in maintaining video quality over extended frames. In this paper, we present a novel, training-free approach for high FPS video generation using pre-trained diffusion models. Our method, DiffuseSlide, introduces a new pipeline that leverages key frames from low FPS videos and applies innovative techniques, including noise re-injection and sliding window latent denoising, to achieve smooth, consistent video outputs without the need for additional fine-tuning. Through extensive experiments, we demonstrate that our approach significantly improves video quality, offering enhanced temporal coherence and spatial fidelity. The proposed method is not only computationally efficient but also adaptable to various video generation tasks, making it ideal for applications such as virtual reality, video games, and high-quality content creation.

视频生成扩散模型高帧率免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。