用掩码引导的物理动力学建模,让物体状态更稳定、预测更准。
MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

- 用掩码对应的图像区域生成位置和动量状态,显式关联物体空间信息
- 在OBJ3D上未来30帧预测,视觉相似性降低25.0%,空间误差下降33.7%
- 适合需要精准物体运动建模的视频预测任务,尤其关注长期稳定性
物体中心的世界模型通过演化一组实体槽来预测未来视频,但通常受动力学监督的变量是无约束的视觉特征。我们提出方法,一种掩码引导的软哈密顿世界模型,使位置类状态明确依赖于槽所拥有的图像支持。一个冻结的视频槽编码器生成槽和掩码;掩码所拥有的支持区域的空间矩形成规范状态 $Q$,时间差分形成 $P$,学习到的能量为有界增量提供软方向偏差。解码器相关的外观和身份信息分别存储在因果视觉上下文中。门控组合器和有界残差将此上下文与传播的相空间状态结合,重构解码器兼容的槽。在OBJ3D上,给定6帧观测并评估后续30帧,该方法相较最强基线,LPIPS降低25.0%,空间MSE降低33.7%。在CLEVRER上,给定6帧观测并评估后续10帧,对应降幅分别为14.5%和18.7%。闭环滚动预测30帧的时序分辨率测量显示,完整模型在整个过程中累积误差更慢。项目页面:https://github.com/moshwm-anon/-moshwm-anon.github.io。
原文摘要 · Abstract (English)
Object-centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We introduce \method{}, a mask-grounded soft-Hamiltonian world model that makes its position-like state explicitly depend on slot-owned image support. A frozen video-slot encoder produces slots and masks; spatial moments of mask-owned support form a canonical state $Q$, temporal differences form $P$, and a learned energy supplies a soft directional bias to a bounded learned increment. Decoder-relevant appearance and identity are stored separately in a causal visual context. A gated composer and bounded residual then combine this context with the propagated phase state to reconstruct decoder-compatible slots. On OBJ3D, given six observed frames and evaluated over the following 30 frames, \method{} reduces LPIPS by 25.0\% and spatial MSE by 33.7\% relative to the strongest object-centric baseline. On CLEVRER, given six observed frames and evaluated over the following ten frames, the corresponding reductions are 14.5\% and 18.7\%. Horizon-resolved visual and object-state measurements show that the complete model accumulates error more slowly throughout the 30-frame closed-loop rollout. Project page:https://github.com/moshwm-anon/-moshwm-anon.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。