让哈密顿视频模型跨时间尺度稳定预测,解决实际应用中的动态泛化难题。
Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models

- 基于连续时间能量函数设计,突破固定步长限制
- 发现非保守场景下两种失效机制并提出针对性修复
- 适合需要多时标模拟的机器人、科学仿真等场景
世界模型通常以固定步长训练预测离散时间物理动态,导致无法在不同时间分辨率下进行预测。这在分层规划、仿真实体迁移及游戏引擎等需多时标查询的应用中尤为关键。哈密顿生成网络(HGN)通过连续时间能量函数提供理论路径,原则上与观测帧率无关。然而,在外力驱动、耗散环境中,其在超出训练范围的时间步长上会失效。我们发现主要失败模式包括:由无约束作用力映射引发的潜在变量幅值增长,以及积分器欠采样导致的全局截断误差累积。针对每种机制提出修正方案,并验证了在远超训练分布的时间分辨率下仍能实现稳定动态预测。详细分析建议若干策略,以实现连续时间视频生成中的时间泛化能力。
原文摘要 · Abstract (English)
World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions. This matters for hierarchical planning, sim-to-real transfer, and scientific or game-engine applications that must query the same dynamics at multiple timescales. Hamiltonian Generative Networks (HGN) offer a principled path forward, grounding predictions in a continuous-time energy function that is, in principle, independent of the observation frame rate. In practice, however, their temporal generalization breaks down in non-conservative settings. We show that in externally forced, dissipative environments, HGN rollouts at step sizes beyond the training regime fail due to distinct failure modes, including latent magnitude growth driven by an unconstrained action-force map, and global truncation error accumulation from an under-resolved integrator. We identify a targeted fix for each mechanism and demonstrate stable dynamics prediction at temporal resolutions well outside the training distribution. In a detailed analysis, we recommend several strategies for enabling temporal generalization in continuous-time video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。