arXiv:2504.14535cs.CV2025-04

通过直接比较生成与真实视频的光流,提升视频扩散模型的运动连贯性。

FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models

  • 直接对比生成与真实视频的光流场,实现运动监督。
  • 在高噪声条件下动态调整损失权重,提升训练稳定性。
  • 适用于需要精准运动控制的机器人视频生成任务。

视频扩散模型(VDMs)能生成高质量视频,但常面临运动不连贯的问题。光学流监督是一种有前景的解决方案,以往工作多采用基于重映射的策略,避免显式匹配光流。本文提出一种新范式——FlowLoss,直接比较生成视频与真实视频中提取的光流场。为应对扩散过程中高噪声导致的光流估计不可靠问题,设计了噪声感知的加权机制,动态调节去噪步骤中的光流损失。在机器人视频数据集上的实验表明,FlowLoss能提升运动稳定性,并加速训练初期的收敛。研究结果为将基于运动的监督引入噪声条件生成模型提供了实用指导。

原文摘要 · Abstract (English)

Video Diffusion Models (VDMs) can generate high-quality videos, but often struggle with producing temporally coherent motion. Optical flow supervision is a promising approach to address this, with prior works commonly employing warping-based strategies that avoid explicit flow matching. In this work, we explore an alternative formulation, FlowLoss, which directly compares flow fields extracted from generated and ground-truth videos. To account for the unreliability of flow estimation under high-noise conditions in diffusion, we propose a noise-aware weighting scheme that modulates the flow loss across denoising steps. Experiments on robotic video datasets suggest that FlowLoss improves motion stability and accelerates convergence in early training stages. Our findings offer practical insights for incorporating motion-based supervision into noise-conditioned generative models.

视频生成扩散模型光流监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。