无需训练即可生成更长视频,解决画面不连贯和重复问题。
FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

- 用滑动窗口+泰迪匹配保持时序一致性和数据流形约束。
- 在长视频生成中显著提升连贯性与画质,支持多倍于原窗口长度的输出。
- 适用于多种模型,可直接扩展至音视频联合生成与文生3D场景。
将视频扩散模型的生成时长扩展至长序列仍是长期存在的重大挑战。现有无需训练的方法分为两类:双向模型的延伸,对特定架构依赖强且随时长增加质量下降;自回归模型则因暴露偏差累积漂移误差,常产生重复运动模式。为此,我们提出一种新颖而简单的推理阶段长视频生成方法,具有架构无关性且无需额外训练。该方法通过重叠滑动窗口生成长视频,利用泰迪匹配融合相邻窗口的预测干净样本,以在重叠区域强制实现流形约束与时间一致性。在高噪声阶段,通过随机早期采样注入新噪声,同步各窗口轨迹;随后切换至确定性微分方程采样,保留精细视觉细节。应用于多种视频生成模型后,本方法生成视频长度可达原始窗口长度的数倍,且在时序连贯性与视觉质量上优于各类无需训练及自回归基线。进一步拓展至音视频联合生成与文本到3DGS,均无需微调。
原文摘要 · Abstract (English)
Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions of bidirectional models, which are tightly coupled to specific architectures and suffer from quality degradation over long horizons, and autoregressive models, which accumulate drift errors due to exposure bias and tend to produce repetitive motion patterns. To address these issues, we propose a novel but simple inference-time approach for long video generation that is architecture-agnostic and requires no additional training. Our method generates long videos via overlapping sliding windows, where predicted clean samples from adjacent windows are blended via \emph{Tweedie matching} to enforce both \textbf{manifold constraint and temporal consistency} across overlap regions. \emph{Stochastic early-phase sampling} then synchronizes per-window trajectories by injecting fresh noise after each Tweedie matching correction in the high-noise phase, before transitioning to deterministic ODE sampling to preserve fine-grained visual fidelity. Applied to various video generation models, our method generates videos several times longer than the native window length while outperforming both training-free and autoregressive baselines in temporal consistency and visual quality, and further extends to audio-video joint generation and text-to-3DGS without any fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。