无需训练,通过分析模型自身预测提升视频生成运动连贯性
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

- 利用模型每步预测的隐空间差异,提取去外观干扰的时间表示
- 通过计算帧间局部方差动态引导采样,降低运动不一致性
- 可直接插入现有模型,提升运动流畅性且不损失画质或提示对齐
文本到视频扩散模型在建模时间特性如运动、物理和动态交互方面存在明显局限。现有方法通过重新训练模型或引入外部条件信号来增强时间一致性。本文探索是否可从预训练模型的预测中直接提取有意义的时间表示,而无需额外训练或辅助输入。提出FlowMo——一种无需训练的引导方法,仅利用模型在每一步扩散过程中的自身预测,增强运动连贯性。FlowMo首先通过测量连续帧对应隐变量之间的距离,构建去外观干扰的时间表示,凸显模型隐含的时间结构;随后通过测量时序维度上的局部方差,估计运动一致性,并在采样过程中动态引导模型降低该方差。在多个文本到视频模型上的大量实验表明,FlowMo显著提升运动连贯性,同时保持视觉质量与提示对齐,为预训练视频扩散模型提供了一种高效即插即用的时间保真增强方案。
原文摘要 · Abstract (English)
Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external conditioning signals to enforce temporal consistency. In this work, we explore whether a meaningful temporal representation can be extracted directly from the predictions of a pre-trained model without any additional training or auxiliary inputs. We introduce FlowMo, a novel training-free guidance method that enhances motion coherence using only the model's own predictions in each diffusion step. FlowMo first derives an appearance-debiased temporal representation by measuring the distance between latents corresponding to consecutive frames. This highlights the implicit temporal structure predicted by the model. It then estimates motion coherence by measuring the patch-wise variance across the temporal dimension and guides the model to reduce this variance dynamically during sampling. Extensive experiments across multiple text-to-video models demonstrate that FlowMo significantly improves motion coherence without sacrificing visual quality or prompt alignment, offering an effective plug-and-play solution for enhancing the temporal fidelity of pre-trained video diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。