通过关系知识迁移,让视频生成更符合物理规律。
Tempered Self-Similarity Alignment for Physically Plausible Video Generation

- 利用时空自相似性捕捉物体交互关系,建模真实动态。
- 在VideoPhy和VideoPhy2上显著提升物理合理性,减少运动失真。
- 适合关注视频生成真实性的研究者与开发者。
尽管视频生成模型取得显著进展,但仍难以生成符合物理规律的视频,常出现外观漂移、不合理运动和时间不一致问题。本文通过将视觉基础模型中的时空自相似性(STSS)所编码的关系知识迁移到视频生成模型中,来解决这一问题。STSS表示空间与时间维度上特征间的成对相似性,揭示了物体在整个视频中与其他实体的交互关系,有效捕捉现实世界中的运动与语义变化。为此,我们提出温度调节的自相似性对齐(TSA)损失,将STSS转化为概率对应分布,并训练视频生成模型在动态区域上对齐该分布。在VideoPhy和VideoPhy2基准上的评估表明,该方法在多种交互场景下显著提升了视频的物理合理性,验证了关系知识迁移在生成真实视频中的有效性。
原文摘要 · Abstract (English)
Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work, we address this limitation by transferring relational knowledge encoded in spatio-temporal self-similarity (STSS) from visual foundation models into video generative models. STSS represents pairwise similarities among features across space and time, revealing the relational structure of how objects interact with other entities throughout a video, effectively capturing real-world dynamics, including object motion and semantic transformations. To transfer this relational knowledge, we propose Tempered Self-similarity Alignment (TSA) loss, which transforms STSS into probabilistic correspondence distributions and trains the video generative model to align its correspondence distributions with those of the visual foundation model on dynamically changing regions. Evaluated on VideoPhy and VideoPhy2 benchmarks, our method demonstrates substantial improvements in physical plausibility across diverse interaction scenarios, validating the effectiveness of transferring relational knowledge for physically realistic video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。