无需训练即可生成更长视频,保持画面连贯与细节清晰
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
- 通过多频段融合,平衡长视频的全局语义与局部细节
- 在4倍和8倍原长视频上显著提升时序一致性与画质
- 可直接接入现有模型,支持多提示与可控生成
近期视频生成模型已在文本驱动短视频生成方面取得高质量进展。然而,将这些模型扩展至更长视频仍面临显著挑战,主要源于时序一致性和视觉保真度下降。初步观察发现,直接将短视频生成模型应用于长序列会导致明显质量退化。进一步分析揭示,随着视频长度增加,高频成分逐渐失真,这一现象称为高频失真。为此,我们提出FreeLong,一种无需训练的框架,在去噪过程中平衡长视频特征的频率分布。FreeLong通过融合全局低频特征(捕捉全视频整体语义)与局部高频特征(从短时间窗口提取,保留细节),实现有效平衡。在此基础上,FreeLong++将双分支设计扩展为多分支架构,包含多个在不同时间尺度运行的注意力分支。通过从全局到局部排列多种窗口尺寸,实现从低频到高频的多带频率融合,确保长视频序列中语义连贯性与精细运动动态。无需额外训练,FreeLong++可无缝集成至现有视频生成模型(如Wan2.1和LTX-Video),生成更长视频,并显著提升时序一致性和视觉保真度。实验表明,该方法在4倍和8倍原长视频生成任务上优于先前方法。同时支持跨提示视频生成、平滑场景切换以及使用长深度或姿态序列进行可控生成。
原文摘要 · Abstract (English)
Recent advances in video generation models have enabled high-quality short video generation from text prompts. However, extending these models to longer videos remains a significant challenge, primarily due to degraded temporal consistency and visual fidelity. Our preliminary observations show that naively applying short-video generation models to longer sequences leads to noticeable quality degradation. Further analysis identifies a systematic trend where high-frequency components become increasingly distorted as video length grows, an issue we term high-frequency distortion. To address this, we propose FreeLong, a training-free framework designed to balance the frequency distribution of long video features during the denoising process. FreeLong achieves this by blending global low-frequency features, which capture holistic semantics across the full video, with local high-frequency features extracted from short temporal windows to preserve fine details. Building on this, FreeLong++ extends FreeLong dual-branch design into a multi-branch architecture with multiple attention branches, each operating at a distinct temporal scale. By arranging multiple window sizes from global to local, FreeLong++ enables multi-band frequency fusion from low to high frequencies, ensuring both semantic continuity and fine-grained motion dynamics across longer video sequences. Without any additional training, FreeLong++ can be plugged into existing video generation models (e.g. Wan2.1 and LTX-Video) to produce longer videos with substantially improved temporal consistency and visual fidelity. We demonstrate that our approach outperforms previous methods on longer video generation tasks (e.g. 4x and 8x of native length). It also supports coherent multi-prompt video generation with smooth scene transitions and enables controllable video generation using long depth or pose sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。