提出结构化世界模型HaM-World,提升长时规划稳定性与鲁棒性。
HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

- 分解潜空间为哈密顿量( q,p)与上下文c,引入选择性记忆建模历史依赖
- 长时规划误差降至基线45%,12项分布外测试中平均收益提升超10%
- 适合需要稳定长程决策的强化学习任务,如机器人控制
世界模型通过学习潜在动力学实现基于模型的规划,但随着规划时长远或动力学分布变化,想象回放会变得不稳定。我们指出,这种不稳定性源于规划器所见潜变量中缺失两种结构:用于近似马尔可夫完备性的历史条件记忆,以及区分构型、动量与任务语义的几何组织。为此,我们提出HaM-World(HMW),将潜状态分解为规范( q, p)子空间和上下文子空间c,同时使用Mamba选择性状态空间记忆作为相同潜动力学的历史条件输入。在此框架内,( q, p)通过能量导出的哈密顿向量场加可学习残差/控制动力学演化,而c捕捉语义、耗散与非保守因素。该设计使规划器共享同一潜状态,用于动力学预测、奖励/价值估计、想象回放与CEM动作搜索。在四个DeepMind Control Suite任务上,HaM-World达到最高平均AUC(117.9,+9.5%),长时回放误差降低至基线的45%,在{3,5,7}个k值的MSE单元中赢下11/12项。面对12种分布外扰动(涵盖动力学漂移、动作延迟、观测掩码),其在每种条件下均取得最高回报,其中Finger Spin平均收益提升10.2%,Reacher Easy达13.6%。机制诊断显示,无动作时哈密顿能量漂移受控,策略回放下能量变化结构清晰,且控制引发的能量转移一致,验证了预期的软哈密顿动力学设计。
原文摘要 · Abstract (English)
World models enable model-based planning through learned latent dynamics, but imagined rollouts become unstable as the planning horizon grows or the dynamics distribution shifts. We argue that this instability reflects two missing structures in planner-facing latents: history-conditioned memory for approximate Markov completeness, and geometric organization that separates configuration, momentum, and task semantics. We propose HaM-World (HMW), a structured world model that decomposes the latent state into a canonical (q, p) subspace and a context subspace c, while using Mamba selective state-space memory as the history-conditioned input to the same latent dynamics. Within this interface, (q, p) evolves through an energy-derived Hamiltonian vector field plus learnable residual/control dynamics, while c captures semantic, dissipative, and non-conservative factors. This gives the planner a single latent state shared by dynamics prediction, reward/value estimation, imagined rollouts, and CEM action search. On four DeepMind Control Suite tasks, HaM-World reaches the highest Avg. AUC (117.9, +9.5%), reduces long-horizon rollout error to 45% of a strong baseline model, and wins 11/12 k in {3,5,7} MSE cells. Under 12 OOD perturbations spanning dynamics shifts, action delay, and observation masking, HaM-World achieves the highest return in every condition, with average OOD-return gains of 10.2% on Finger Spin and 13.6% on Reacher Easy. Mechanism diagnostics further show bounded action-free Hamiltonian-energy drift, structured energy variation under policy rollouts, and coherent control-induced energy transfer, supporting the intended Soft-Hamiltonian dynamics design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。