arXiv:2608.05976cs.CV2026-08中稿 · publication in ACM…被引 1

无需训练即可生成高质量长视频,解决时序连贯性与动作多样性难题。

Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

论文配图:Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
图 1 · 摘自论文原文
  • 通过三种互补策略实现无训练长视频生成
  • 在VBench-Long上优于现有基线模型的时序连贯性与动作多样性
  • 适用于不同结构的视频扩散模型,适合需要快速部署长视频生成的场景

近期扩散模型在视频生成领域取得显著进展,但多数现有模型基于短视频训练,外推至长视频时难以保持长时序一致性并保留丰富动作。为生成一致、高质量且动态的长视频,我们提出Diff-VF——一种无需训练、即插即用且模型无关的框架,可在不修改或微调基础模型的前提下,将现有短视频扩散模型转化为长视频生成器。Diff-VF结合三种互补策略:混合噪声初始化(HNI)约束全局语义,加权窗口采样(WWS)消除窗口间不连续性,时间扩展采样(TES)通过时步可变融合建立长程依赖。进一步引入跳过残差引导(Skip Residual Guidance)实现长视频增强,通过时步依赖引导平衡保真度与真实感。VBench-Long评估显示,Diff-VF在时序连贯性与动作多样性之间取得更优平衡,优于基线模型及近期无训练方法(FreeNoise、FreeLong、RIFLEx),同时保持竞争性的帧级质量。在两种基础模型上的实验验证其对不同时空建模策略的适用性。大量消融实验确认各组件与超参数的有效性。

原文摘要 · Abstract (English)

Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range temporal coherence while retaining diverse motions. To generate consistent, high-quality and dynamic long videos, we propose Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model. Diff-VF couples three complementary strategies: Hybrid Noise Initialization (HNI) to constrain global semantics, Weighted Window Sampling (WWS) to remove inter-window discontinuities, and Temporal Extended Sampling (TES) to establish long-range dependencies with a timestep-varying fusion. We further extend Diff-VF to long-video enhancement via Skip Residual Guidance that balances fidelity and realism through timestep-dependent guidance. VBench-Long evaluation results show that Diff-VF achieves a more favorable balance between temporal coherence and motion diversity than base models and recent training-free long video generation baselines, including FreeNoise, FreeLong, and RIFLEx, while maintaining competitive frame-wise quality. Experiments on two base models demonstrate the applicability to video diffusion models with different spatial-temporal modeling strategies. Extensive ablations validate the contribution of each component and hyperparameters.

视频生成扩散模型长视频无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。