用最小作用量原理让视觉世界模型更符合物理规律,提升长时预测准确性。
LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations

- 在视觉隐空间中构建基于最小作用量的离散变分积分器,使未来预测遵循物理规律。
- 在合成物理和机器人交互任务中,相比基线模型,运动更平滑、背景更一致、误差更小。
- 适合需要长期准确预测的机器人规划与具身智能系统,尤其关注物理一致性场景。
从视觉观测中学习预测性世界模型是具身人工智能的核心问题,广泛应用于基于模型的强化学习和机器人规划。现有潜在世界模型通常使用无约束的神经转移函数生成未来状态,而现代视频生成系统虽强调感知合理性,但常依赖辅助损失、外部引导或独立动力学模块引入物理结构。导致长时程滚动预测仍缺乏对真实动力学物理原则的有效约束,出现累积误差、能量漂移及物理不一致的未来。本文提出最小作用量世界模型(LaWM),将最小作用量原理应用于学习的视觉隐空间:未来轨迹由学习的拉格朗日作用泛函决定,而非仅依赖无约束的转移预测器。核心技术实现为潜在变分积分器:LaWM将观测编码为学习的广义坐标,学习连续隐状态间的潜在离散拉格朗日量,构建离散作用泛函,并通过求解相应离散积分条件推进预测。因此,物理结构不仅用于评估、正则化或约束已生成轨迹,而是直接定义了潜在转移规则本身。由于转移由离散变分原理驱动,LaWM为长时程视觉预测提供了结构保持偏差。在物理纯净的合成动态和具身机器人交互基准上,LaWM在物理不变性、背景一致性、运动平滑性以及外观与几何预测指标上均优于视频生成与世界模型基线。
原文摘要 · Abstract (English)
Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning and robotic planning. Existing latent world models typically generate future states with unconstrained neural transition functions, while modern video generation systems often prioritize perceptual plausibility or introduce physical structure through auxiliary losses, external guidance, or separate dynamics modules. As a result, long-horizon rollouts can remain weakly grounded in the physical principles that govern real dynamics, leading to compounding error, energy drift, and physically inconsistent futures. We propose Least Action World Models (LaWM), a latent world-modeling framework that operationalizes the Principle of Least Action in learned visual latent space: future rollouts are governed by a learned Lagrangian action functional rather than produced only by an unconstrained transition predictor. Our main technical realization is a latent variational integrator: LaWM encodes observations into learned generalized coordinates, learns a latent discrete Lagrangian over consecutive latent states, constructs a discrete action functional, and advances prediction by solving the corresponding discrete integration condition. Thus, physical structure is not merely used to score, regularize, or constrain a completed trajectory; it defines the latent transition rule itself. Because the transition is induced by a discrete variational principle, LaWM provides a structure-preserving bias for long-horizon visual prediction. Across physics-clean synthetic dynamics and embodied robot interaction benchmarks, LaWM improves physical invariance, background consistency, motion smoothness, and appearance and geometric prediction metrics over video-generation and world-model baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。