让机器人模型学会预测物体惯性与环境隐流的影响
FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift

- 分解视觉历史为运动状态和环境漂移上下文
- 在模拟水域车辆中,长时预测误差降低37%
- 适合水下/风力等有惯性与隐流的场景
机器人学习中的世界模型从视觉观测和动作中预测未来状态,使智能体能够推理控制后果。然而,多数动作条件模型在以即时控制为主导的环境中评估,而真实世界物体(如水面航行器)受惯性影响持续运动,并受水流动、风等不可见环境漂移影响。我们提出 FlowMo-WM,一种端到端可训练的视觉世界模型,无需流场监督,即可从图像-动作历史中推断出物体中心的运动状态和关联的长期隐藏漂移上下文。该模型将图像-动作历史分解为短历史潜在状态(表征物体运动)和长历史上下文(表征缓慢变化的外部影响),并通过零上下文残差转移,在潜在空间演进中分离动作主导的动力学与上下文依赖的漂移效应。在包含多种隐藏流、扰动和随机车辆动力学的模拟水面车辆环境中,FlowMo-WM 在长时程预测上优于代表性动作条件潜伏世界模型。预测时上下文消融实验表明,推断的环境上下文对隐藏漂移下的稳定预测至关重要;冻结线性探测器则揭示了学习因子所编码的信息。
原文摘要 · Abstract (English)
World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-conditioned models are evaluated in settings where motion is dominated by immediate control, whereas aquatic surface vehicles and other real-world objects continue moving under inertia and are displaced by hidden ambient drift, such as water currents or wind. We propose FlowMo-WM, an end-to-end trainable visual world model that infers object-centric motion state and a predictive long-history context associated with hidden drift from image-action histories without direct supervision of flow fields. FlowMo-WM factorizes image-action history into a short-history latent state, trained to summarize object-centric motion, and a longer-history context, trained to summarize slowly varying exogenous influences. A zero-context residual transition separates action-conditioned base dynamics from context-dependent drift effects during latent rollout. In simulated aquatic surface-vehicle environments with diverse hidden flows, disturbances, and randomized vehicle dynamics, FlowMo-WM improves long-horizon rollout accuracy over representative action-conditioned latent world models. Prediction-time context ablations, in which the inferred context is zeroed or shuffled during rollout, show that the ambient context is important for stable prediction under hidden drift, while frozen linear probes characterize information encoded in the learned factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。