提出频域前瞻修正方法,提升文本生成视频的运动连贯性。
SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

- 通过频域前瞻预测干净潜在表示并修正幅度谱
- 在Wan2.2上显著减少物理伪影,提升运动一致性
- 仅增加4次额外采样,适合部署于现有生成框架
流匹配通过潜在常微分方程采样实现了鲁棒的文本到视频生成。然而,速度近似和数值离散化误差不可避免地累积,导致采样轨迹漂移,使生成视频出现严重的时空不一致。直接纠正这些漂移的噪声潜在表示具有挑战性:(i) 时间步依赖的噪声掩盖了可靠的结构线索;(ii) 空间干预可能破坏复杂局部几何,且计算开销大。为此,我们提出频域前瞻修正(SpecLoR),一种即插即用的推理方法,通过前瞻预测避开噪声,并通过将修正转移到频域来规避时空耦合,该域中自然视频的通用统计先验易于获取。首先,在早期采样阶段,SpecLoR前瞻估计干净潜在表示 $z_{t,0}$ 并计算其三维时空谱。接着,修正幅度谱以匹配先验,保留相位不变。最后,将修正状态重新加噪以恢复常微分方程积分。在 Wan2.2 上的实验表明,SpecLoR 在多个基准上显著降低物理伪影并提升运动连贯性,仅增加 4 次额外的 NFE(NFE: Number of Function Evaluations)。
原文摘要 · Abstract (English)
Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors inevitably accumulate, causing sampling trajectories to drift. Consequently, generated videos often suffer from severe spatiotemporal inconsistencies. Nevertheless, directly correcting these drifted, noisy latents is challenging: (i) timestep-dependent noise obscures reliable structural cues; (ii) spatial interventions risk disrupting intricate local geometry while incurring heavy computational costs. To address this, we propose Spectral Lookahead Rectification (SpecLoR), a plug-and-play inference method that bypasses noise via lookahead prediction, and circumvents spatiotemporal entanglement by shifting corrections to the frequency domain, where universal statistical priors of natural videos are readily available. First, during early sampling stages, SpecLoR looks ahead to estimate the clean latent $z_{t,0}$ and computes its 3D spatiotemporal spectrum. Next, SpecLoR rectifies the amplitude spectrum to match the prior, leaving the phase intact. Finally, the corrected state is re-noised to resume ODE integration. Experiments on Wan2.2 demonstrate that SpecLoR significantly reduces physical artifacts and enhances motion coherence across multiple benchmarks with minimal computational overhead (4 additional NFEs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。