无需微调,通过同步采样实现多事件长视频生成的连贯性
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
- 通过同步反向与优化采样,对齐去噪路径以保证长视频一致性
- 在多个提示下生成长视频,相比现有方法过渡更平滑、语义更连贯
- 适合需要高一致性的长视频生成场景,如影视创作或智能编辑
尽管近期文本到视频的扩散模型能在单个提示下生成高质量短视频,但一次性生成真实世界中的长视频仍面临数据有限和计算成本高的挑战。为解决此问题,已有工作提出无需微调的方法,即通过多个提示扩展现有模型以实现动态可控的内容变化。然而,这些方法主要关注相邻帧间的平滑过渡,常导致内容漂移和长序列中语义连贯性的逐渐丧失。为此,本文提出同步耦合采样(SynCoS),一种新型推理框架,通过同步整个视频的去噪路径,确保远距离帧间也保持长期一致性。该方法结合反向采样与基于优化的采样:前者保障局部过渡流畅,后者维持全局语义一致。直接交替使用两者会因去噪轨迹错位而破坏提示引导,引入非预期内容变化。SynCoS通过固定时间步和基准噪声实现同步,确保完全耦合的采样过程。大量实验表明,SynCoS显著提升多事件长视频生成效果,在定量与定性评估上均优于先前方法。
原文摘要 · Abstract (English)
While recent advancements in text-to-video diffusion models enable high-quality short video generation from a single prompt, generating real-world long videos in a single pass remains challenging due to limited data and high computational costs. To address this, several works propose tuning-free approaches, i.e., extending existing models for long video generation, specifically using multiple prompts to allow for dynamic and controlled content changes. However, these methods primarily focus on ensuring smooth transitions between adjacent frames, often leading to content drift and a gradual loss of semantic coherence over longer sequences. To tackle such an issue, we propose Synchronized Coupled Sampling (SynCoS), a novel inference framework that synchronizes denoising paths across the entire video, ensuring long-range consistency across both adjacent and distant frames. Our approach combines two complementary sampling strategies: reverse and optimization-based sampling, which ensure seamless local transitions and enforce global coherence, respectively. However, directly alternating between these samplings misaligns denoising trajectories, disrupting prompt guidance and introducing unintended content changes as they operate independently. To resolve this, SynCoS synchronizes them through a grounded timestep and a fixed baseline noise, ensuring fully coupled sampling with aligned denoising paths. Extensive experiments show that SynCoS significantly improves multi-event long video generation, achieving smoother transitions and superior long-range coherence, outperforming previous approaches both quantitatively and qualitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。