arXiv:2605.06509cs.CV2026-05被引 1

无需训练即可生成长视频,解决动态模糊与时间不一致问题

FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction

论文配图:FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction
图 1 · 摘自论文原文
  • 从奇异谱视角重构视频,分离全局低秩结构与局部高秩细节
  • 在Wan2.1和LTX-Video上显著提升长视频时间动态表现
  • 适合需要高质量长视频生成且无训练成本的场景

视频扩散模型在短视频生成中表现良好,但其无训练扩展至长视频时常出现内容漂移、时间不一致和动态过平滑问题。现有方法通过全局分支与局部分支结合提升时间一致性,但通常依赖预设规则将外观一致性与时间动态解耦,当外观与动作进展紧密耦合(如镜头运动、序列动作)时,该划分不可靠。本文从奇异谱角度分析视频时序扩展问题,发现扩大自注意力窗口会导致谱集中:能量集中于少数低秩奇异方向,保留粗略结构但抑制高秩空间细节与丰富时间变化。为此,提出FreeSpec——一种无训练的谱重构框架。它对全局与局部特征进行奇异值分解,以全局分支作为低秩谱引导,局部分支作为高秩重建基础,实现谱级融合,避免了先前刚性特征划分,既保持长程一致性,又更好保留空间细节与时间动态。在Wan2.1与LTX-Video上的实验表明,FreeSpec显著提升长视频生成效果,尤其在时间动态方面,同时维持强视觉质量与时间一致性。

原文摘要 · Abstract (English)

Video diffusion models perform well in short-video synthesis, but their training-free extension to long videos often suffers from content drift, temporal inconsistency, and over-smoothed dynamics. Existing methods improve temporal consistency by combining a global branch with a local branch, but they often further decompose appearance consistency and temporal dynamics within each branch using predefined criteria. This assignment is unreliable when appearance and action progression are tightly coupled, such as in camera motion and sequential motion. We analyze the video temporal extension issue from a singular-spectrum perspective and show that enlarged self-attention windows induce spectral concentration: spectral energy becomes dominated by a few low-rank singular directions, preserving coarse structure but suppressing high-rank spatial details and motion-rich temporal variations. To mitigate this problem, we propose FreeSpec, a training-free spectral reconstruction framework for long-video generation. FreeSpec decomposes global and local features with singular value decomposition, and uses the global branch as low-rank spectral guidance and the local branch as a high-rank reconstruction basis. This spectrum-level fusion avoids the rigid feature partitioning of previous decomposition rules, preserving long-range consistency while better retaining spatial details and temporal dynamics. Experiments on Wan2.1 and LTX-Video demonstrate that FreeSpec improves long-video generation, especially for temporal dynamics, while maintaining strong visual quality and temporal consistency. Project demo: https://fdchen24.github.io/FreeSpec-Website/.

视频生成扩散模型长视频无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。