arXiv:2501.09019cs.CV2025-01AAAI被引 14

无需微调即可生成任意长度且内容一致的长视频,解决帧间一致性难题。

Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion

论文配图:Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
图 1 · 摘自论文原文
  • 引入新型潜空间采样与跨帧注意力机制,提升帧间结构与主体一致性。
  • 在VBench上生成120秒视频,主体一致性和运动流畅性显著优于基线。
  • 适合需要长时序连贯性的视频生成场景,如影视创作、虚拟人直播。

基于预训练文本到视频模型的先进方法——先进先出(FIFO)视频扩散,已成为无需微调生成长视频的有效手段。该方法通过维护一个带逐渐增加噪声的帧队列,在队列头部持续生成清晰帧,同时在尾部添加高斯噪声。然而,由于缺乏帧间的对应建模,FIFO-Diffusion常难以保持长时序的一致性。本文提出Ouroboros-Diffusion,一种新颖的视频去噪框架,旨在增强结构与内容(主体)一致性,实现任意长度视频的稳定生成。具体而言,我们在队列尾部引入新的潜空间采样技术以改善结构一致性,确保帧间过渡自然;设计了主体感知跨帧注意力(SACFA)机制,在短片段内对齐主体,提升视觉连贯性;此外,引入自循环引导机制,利用队列前端所有较清晰帧的信息,指导尾端噪声帧的去噪过程,促进全局上下文信息的丰富交互。在VBench基准上的大量实验表明,我们的Ouroboros-Diffusion在主体一致性、运动平滑性和时间一致性方面均表现卓越。

原文摘要 · Abstract (English)

The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation. This technique maintains a queue of video frames with progressively increasing noise, continuously producing clean frames at the queue's head while Gaussian noise is enqueued at the tail. However, FIFO-Diffusion often struggles to keep long-range temporal consistency in the generated videos due to the lack of correspondence modeling across frames. In this paper, we propose Ouroboros-Diffusion, a novel video denoising framework designed to enhance structural and content (subject) consistency, enabling the generation of consistent videos of arbitrary length. Specifically, we introduce a new latent sampling technique at the queue tail to improve structural consistency, ensuring perceptually smooth transitions among frames. To enhance subject consistency, we devise a Subject-Aware Cross-Frame Attention (SACFA) mechanism, which aligns subjects across frames within short segments to achieve better visual coherence. Furthermore, we introduce self-recurrent guidance. This technique leverages information from all previous cleaner frames at the front of the queue to guide the denoising of noisier frames at the end, fostering rich and contextual global information interaction. Extensive experiments of long video generation on the VBench benchmark demonstrate the superiority of our Ouroboros-Diffusion, particularly in terms of subject consistency, motion smoothness, and temporal consistency.

视频生成扩散模型长视频一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。