用隐空间动态差异引导损失权重,提升视频生成的运动质量。
Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
- 引入隐空间帧间差异作为运动先验,动态调整损失权重。
- 在VBench和VMBench上分别提升3.31%和3.58%,显著改善运动细节。
- 适合需要高动态保真度的视频生成任务,如动作密集场景。
视频生成模型在静态场景中已取得显著进展,但在剧烈动态变化的视频生成中表现受限,质量因噪声破坏时序一致性而下降,且难以学习动态区域。现有扩散模型对所有场景使用静态损失,限制了其捕捉复杂动态的能力。为此,本文提出将隐空间中的帧间差异(Latent Temporal Discrepancy, LTD)作为运动先验,用于指导损失加权:对隐空间差异大的区域施加更大惩罚,稳定区域则维持常规优化。该策略提升了训练稳定性,增强了对高频动态的重建能力。在通用基准VBench和运动聚焦基准VMBench上的实验表明,本方法分别优于强基线3.31%和3.58%,显著提升运动质量。
原文摘要 · Abstract (English)
Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic regions. {Unfortunately, existing diffusion models rely on static loss for all scenarios, constraining their ability to capture complex dynamics.} To address this issue, we introduce Latent Temporal Discrepancy (LTD) as a motion prior to guide loss weighting. LTD measures frame-to-frame variation in the latent space, assigning larger penalties to regions with higher discrepancy while maintaining regular optimization for stable regions. This motion-aware strategy stabilizes training and enables the model to better reconstruct high-frequency dynamics. Extensive experiments on the general benchmark VBench and the motion-focused VMBench show consistent gains, with our method outperforming strong baselines by 3.31% on VBench and 3.58% on VMBench, achieving significant improvements in motion quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。