arXiv:2605.23878cs.CV2026-05被引 3

从无标签视频中提取运动先验,提升视频生成的物理真实感。

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

论文配图:LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
图 1 · 摘自论文原文
  • 自监督学习视频帧间隐空间运动变化,构建运动先验
  • 在两个物理一致性数据集上超越有外部监督的基线模型
  • 可无缝集成到现有扩散模型,无需修改架构或输入输出

当前视频生成模型虽视觉效果出色,但在物理和运动一致性方面仍存在不足,限制其作为可靠世界模拟器的应用。现有方法多依赖外部模拟器、教师模型或专门的物理数据集。本文探索一种互补的自监督路径:从训练视频扩散模型所用的未标注视频中提取运动线索。提出LaMo,通过条件于当前隐状态和提示词,对帧间隐空间变化建立运动先验。该先验通过两个轻量级模块实现:训练时使用宏观运动漂移作为运动漂移损失,采样时使用学习的微观运动场作为运动先验引导。两者均与现有视频扩散主干模型即插即用,无需架构或输入输出改动。在VideoPhy和VideoPhy2数据集上,LaMo提升了CogVideoX模型表现,并优于近期依赖外部监督的物理感知基线。在VBench评测中,保持整体生成质量的同时显著改善了运动相关维度。结果表明,未标注视频中蕴含可用于提升现代视频扩散模型物理保真度的有效运动监督。

原文摘要 · Abstract (English)

Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedies often rely on external simulators, teacher models, or curated physics-focused data. We explore a complementary self-supervised direction: extracting motion cues from the unlabeled videos already used to train video diffusion models. We propose LaMo, which formulates a latent motion prior over frame-to-frame latent changes conditioned on the current latent and prompt. This prior is exposed through two lightweight readouts: a macro motion drift used during training as a Motion Drift Loss, and a learned micro motion field used during sampling as Motion Prior Guidance. Both components are plug-and-play with existing video diffusion backbones, requiring no architectural or I/O changes. On VideoPhy and VideoPhy2, LaMo improves CogVideoX backbones and outperforms recent physics-aware baselines that use external supervision. On VBench, it preserves overall generation quality while improving motion-related dimensions. These results suggest that unlabeled video contains useful motion supervision for improving physical fidelity in modern video diffusion models.

视频生成扩散模型物理一致性自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。