arXiv:2608.10780cs.RO2026-08被引 1

提出分阶段预测框架,让机器人更懂任务进展

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

论文配图:StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
图 1 · 摘自论文原文
  • 用联合嵌入预测架构分步预测任务阶段进展
  • 在50个任务中达成90.25%成功率,减少近6%执行步数
  • 适合需要理解任务流程的通用机器人控制场景

通用机器人策略需将多模态观测和语言指令映射到多样化任务的动作。现有方法通常将未来表示为固定、短期的视频-动作片段,仅捕捉局部场景演变,未显式描述从当前阶段到下一阶段的任务推进过程。为此,我们区分两种互补的未来:短期物理未来用于捕捉局部场景变化,阶段级语义未来用于表示任务进展。提出StageWAM,基于Motus的World Action Model(WAM)引入目标条件的联合嵌入预测架构(Stage-JEPA)。给定当前观测和任务指令,Stage-JEPA利用冻结的V-JEPA2编码器提取当前状态表示,并预测下一推断阶段的潜在目标。在50个RoboTwin 2.0任务(含清洁与随机化环境)上,StageWAM实现90.25%的整体成功率,成功轨迹的平均执行步数相较最强基线减少5.97%。

原文摘要 · Abstract (English)

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

机器人操控任务阶段预测世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。