arXiv:2605.10942cs.RO2026-05被引 6

提出HarmoWAM,让机器人同时实现泛化移动与精准操作。

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

论文配图:HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
图 1 · 摘自论文原文
  • 用世界模型驱动预测与反应双专家协同
  • 零样本测试中性能超越现有模型33%以上
  • 适合需要灵活适应新环境的机器人任务

世界动作模型(WAMs)通过建模物理动态成为机器人控制的有前景范式。当前方法主要分为两类:‘想象后执行’(基于视频预测推断动作)和‘联合建模’(联合建模动作与视觉表示)。系统实验发现两者存在根本权衡:前者虽具泛化迁移能力但交互精度不足,后者虽能生成精细时序动作却受限于训练分布探索空间。为此,我们提出端到端的HarmoWAM,利用世界模型统一预测与响应控制,实现泛化移动与精确操作。具体而言,世界模型提供时空物理先验,分别指导两个互补的动作专家:一个利用隐空间动态迭代生成动作的预测专家,另一个直接从预测视觉演化中推断动作的反应专家。为实现自适应协调,引入过程自适应门控机制,自动决定切换时机与位置。该机制使世界模型可驱动反应专家拓展探索空间,同时让预测专家在任务不同阶段完成精准交互。评估方面,我们在六个真实机器人任务上构建了三个训练未见测试环境,涵盖背景、位置及物体语义变化。结果表明,HarmoWAM在这些场景中表现出强零样本泛化能力,显著优于先前最优的VLA模型和WAMs,提升幅度分别为33%和29%。

原文摘要 · Abstract (English)

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Execute" approach, which uses video prediction to infer actions via inverse dynamics, and the "Joint Modeling" approach, which jointly models actions and video representations. Based on systematic experiments, we observe a fundamental trade-off between these paradigms: the former explicitly leverages world models for generalizable transit but lacks interaction precision, whereas the latter enables fine-grained, temporally coherent action generation but is constrained by the exploration space of the training distribution. Motivated by these findings, we propose HarmoWAM, an end-to-end WAM that fully leverages a world model to unify predictive and reactive control, enabling both generalizable transit and precise manipulation. Specifically, the world model provides spatio-temporal physical priors that condition two complementary action experts: a predictive expert that leverages latent dynamics for iterative action generation, and a reactive expert that directly infers actions from predicted visual evolution. To enable adaptive coordination, a Process-Adaptive Gating Mechanism is proposed to automatically determine the timing and location of switching between them. This allows the world model to drive the reactive expert to expand the exploration space and the predictive expert to perform precise interactions across different stages of a task. For evaluation, we construct three training-unseen test environments across six real-world robotic tasks, covering variations in background, position, and object semantics. Notably, HarmoWAM achieves strong zero-shot generalization across these scenarios, significantly outperforming prior state-of-the-art VLA models and WAMs by margins of 33% and 29%, respectively.

机器人控制世界模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。