arXiv:2606.13674cs.CV2026-06被引 5

用语义视觉动作标记构建世界动作模型,提升机器人指令跟随能力。

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

论文配图:RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
图 1 · 摘自论文原文
  • 设计语义对齐的视觉-动作标记化器,将视觉输入转为联合表征
  • 在真实与仿真任务中实现强泛化性能,优于基于像素重建的方法
  • 适合追求通用机器人策略的研究者和开发者

本文提出RepWAM,一种基于语义视觉-动作标记化器的表征中心型世界动作模型(WAM)。现有WAM通常沿用预训练视频生成模型中的重建导向视频标记化器,虽能保持视觉保真度,但仅依赖像素重建难以指导学习连接未来预测与机器人控制的指令跟随动态。为此,我们探索了一种语义视觉-动作潜在空间,训练一个将视觉输入映射为对齐的视觉与潜在动作标记的表示视觉-动作标记化器。随后,预训练WAM以联合建模语言指令下的未来视觉状态及连接它们的潜在动作,并适配真实机器人轨迹实现闭环操作。在真实世界操作任务与仿真基准上的实验表明,RepWAM在多种操作场景中表现优异,消融实验进一步验证了语义视觉-动作标记化相比重建导向方法的价值。结果确立了该标记化方式作为世界动作模型的有力基础,推动通用机器人策略的发展。代码与权重将公开于https://github.com/wdrink/RepWAM。

原文摘要 · Abstract (English)

This work presents RepWAM, a representation-centric world action model (WAM) built on representation visual-action tokenizers. Existing WAMs typically inherit reconstruction-oriented video tokenizers from pretrained video generation models. Although these tokenizers preserve visual fidelity, pixel reconstruction alone provides limited guidance for learning instruction-following dynamics that connect future prediction with robot control. To address this, we explore a semantic visual-action latent space for representation-centric world action modeling. Specifically, we train a representation visual-action tokenizer that maps visual inputs into aligned visual and latent action tokens. We then pretrain our WAM to jointly model future visual states and the latent actions that connect them under language instructions, followed by adaptation to real robot trajectories for closed-loop manipulation. Experiments on real-world manipulation tasks and simulation benchmarks show that RepWAM delivers strong performance across diverse manipulation settings, while ablations highlight the value of semantic visual-action tokenization over reconstruction-oriented alternatives. These results establish representation visual-action tokenization as a promising foundation for world action models and a step toward generalist robot policies. Code and weights will be available at https://github.com/wdrink/RepWAM.

世界动作模型机器人控制语义标记化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。