arXiv:2501.05484cs.CV2025-01

无需微调即可生成高质量长视频,解决时序不一致问题

Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion

  • 通过全局-局部协同去噪构建稳定生成轨迹
  • 在3倍和6倍长度视频上保持内容连贯性与视觉质量
  • 适合希望直接使用现有模型生成长视频的研究者

生成高保真、连贯的长视频是当前研究的重要目标。尽管近期视频扩散模型展现出潜力,但仍面临时空不一致和计算资源消耗高的问题。本文提出GLC-Diffusion,一种无需微调的长视频生成方法。该方法通过建立全局-局部协同去噪路径,确保整体内容一致性与帧间时序连贯性。此外,引入噪声重初始化策略,结合局部噪声打乱与频域融合,提升全局内容一致性和视觉多样性。进一步提出视频运动一致性精修(VMCR)模块,通过像素级与频域损失梯度优化,增强视觉一致性和时间平滑性。大量实验,包括在不同长度视频(如3倍和6倍)上的定量与定性评估,表明该方法能有效集成于现有视频扩散模型,生成优于以往方法的连贯、高保真长视频。

原文摘要 · Abstract (English)

Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational resource demands. We propose GLC-Diffusion, a tuning-free method for long video generation. It models the long video denoising process by establishing denoising trajectories through Global-Local Collaborative Denoising to ensure overall content consistency and temporal coherence between frames. Additionally, we introduce a Noise Reinitialization strategy which combines local noise shuffling with frequency fusion to improve global content consistency and visual diversity. Further, we propose a Video Motion Consistency Refinement (VMCR) module that computes the gradient of pixel-wise and frequency-wise losses to enhance visual consistency and temporal smoothness. Extensive experiments, including quantitative and qualitative evaluations on videos of varying lengths (\textit{e.g.}, 3\times and 6\times longer), demonstrate that our method effectively integrates with existing video diffusion models, producing coherent, high-fidelity long videos superior to previous approaches.

视频生成扩散模型长视频去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。