arXiv:2501.02741cs.CV2025-01被引 4

无需训练即可生成任意长度高质量长视频。

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

  • 分段逐块去噪,模拟砌墙过程增强帧间联系。
  • 生成视频质量显著优于现有方法,无明显退化。
  • 适合需要长视频生成且无训练资源的场景。

扩散模型在文本驱动视频生成方面取得显著进展。然而,训练长视频生成模型需大量计算资源和数据,导致多数视频扩散模型仅限于少量帧。现有无需训练的方法利用预训练短视频扩散模型生成长视频时,常面临运动动态不足、视频保真度下降等问题。本文提出 Brick-Diffusion,一种新颖的无需训练方法,可生成任意长度的长视频。该方法引入砖块到墙体去噪策略:在潜空间中分段去噪,并在后续迭代中施加步长。这一过程模拟错缝砌墙结构,每块砖代表一个去噪片段,促进帧间信息交互,提升整体视频质量。定量与定性评估表明,Brick-Diffusion 在生成高保真视频方面优于现有基线方法。

原文摘要 · Abstract (English)

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos.

视频生成扩散模型长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。