提出新型移动操作世界模型,提升长程任务成功率与精细控制精度
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

- 引入中间隐动作,对齐时间粒度与机器人控制
- 双层混合变换器解耦移动与操作动作空间,减少冲突
- 梦想强迫训练增强推理一致性,适合复杂长程任务研究者
移动操作是通用机器人的关键能力,但现有具身学习方法仍面临挑战。视觉语言动作(VLA)策略多为反应式且缺乏显式世界建模,而现有世界动作模型(WAM)在时间粒度、动作空间和训练-测试一致性方面与移动操作结构不匹配:它们基于粗粒度视频片段,建模纠缠的导航-操作动作,并在与自回归推理不符的监督下训练逆动力学,导致忽略细粒度接触动态、动作分布冲突,并在长程推演中累积误差。本文提出ABot-M0.5,基于三重对齐理念——时间粒度、动作空间与训练-测试一致性。为对齐时间粒度,引入中间隐动作以捕捉局部视觉状态变化,作为视频隐表示与具身控制之间的桥梁;为对齐动作空间,设计双层混合变压器架构,解耦模态表征与异质动作子空间(如底盘移动与机械臂操作);为对齐推理条件,提出梦想强迫训练策略,在模型预测视频上逐步训练逆动力学,提升训练-测试一致性与自回归预测鲁棒性。在具有挑战性的移动与精细操作基准上,ABot-M0.5在长程任务成功率与精细控制精度上均达到当前最优表现。结果凸显了粒度对齐、动作解耦与推理一致性的关键作用。
原文摘要 · Abstract (English)
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。