arXiv:2601.03665cs.CV2026-01被引 3

让视频生成更符合物理规律,通过潜空间引导实现真实动态效果。

PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance

  • 从扩散潜变量中预测物理特征并注入生成器注意力层
  • 在Latte模型上实现稳定训练,恢复出预训练的物理表征
  • 适合关注物理一致性与生成质量的视频生成研究者

当前视频生成模型虽能产出高质量视觉内容,但常缺乏对真实物理动态的建模能力,导致物体碰撞不自然、重力异常及时间闪烁等伪影。本文提出PhysVideoGenerator,一种概念验证框架,将可学习的物理先验显式嵌入生成过程。设计轻量级预测网络PredictorP,直接从噪声扩散潜变量中回归预训练视频联合嵌入预测架构(V-JEPA 2)提取的高层物理特征。这些预测的物理标记通过专用交叉注意力机制注入基于DiT的生成器(Latte)的时序注意力层。主要贡献在于验证了该联合训练范式的可行性:证明扩散潜变量包含足够信息以恢复V-JEPA 2的物理表征,且多任务优化在训练过程中保持稳定。本报告详述了架构设计、技术挑战与训练稳定性验证,为未来大规模物理感知生成模型评估奠定基础。

原文摘要 · Abstract (English)

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unnatural object collisions, inconsistent gravity, and temporal flickering. In this work, we propose PhysVideoGenerator, a proof-of-concept framework that explicitly embeds a learnable physics prior into the video generation process. We introduce a lightweight predictor network, PredictorP, which regresses high-level physical features extracted from a pre-trained Video Joint Embedding Predictive Architecture (V-JEPA 2) directly from noisy diffusion latents. These predicted physics tokens are injected into the temporal attention layers of a DiT-based generator (Latte) via a dedicated cross-attention mechanism. Our primary contribution is demonstrating the technical feasibility of this joint training paradigm: we show that diffusion latents contain sufficient information to recover V-JEPA 2 physical representations, and that multi-task optimization remains stable over training. This report documents the architectural design, technical challenges, and validation of training stability, establishing a foundation for future large-scale evaluation of physics-aware generative models.

视频生成物理建模扩散模型潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。