arXiv:2607.23969cs.RO2026-07

用隐空间预测替代视觉生成,让机器人模型更高效更鲁棒。

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

论文配图:LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
图 1 · 摘自论文原文
  • 以预测性语义对齐为核心,直接在隐空间建模物理动态。
  • 在LIBERO和RoboTwin 2.0上达顶尖性能,无需大规模预训练。
  • 支持零开销推理与真实世界迁移,适合工业级机器人部署。

世界动作模型(WAMs)是具身智能的重要范式,但传统依赖像素级视频生成的方法存在根本瓶颈:重建无关视觉细节会耗散表征能力,使策略易受视觉干扰。本文提出LeapBot-WA,通过将联合嵌入预测架构(JEPA)作为世界锚点,建立全新的预测-隐空间范式。该模型摒弃视觉合成,转而聚焦于预测语义对齐,在隐空间中直接提取抽象物理动态。为弥合非高斯预测特征与扩散先验之间的模态差异,提出各向同性语义自编码器(ISAE),将锚点隐空间重构为适合扩散的流形,防止离流形漂移。同时设计非对称混合变压器(MoT)结构:训练时由动力学专家分支引导动作扩散分支;推理时裁剪重型动力学分支,实现零开销执行。LeapBot-WA在LIBERO上达到当前最优预测模型性能,在RoboTwin 2.0上媲美顶级生成式WAMs,且无需大规模轨迹预训练。进一步展现出对未见环境的优异零样本鲁棒性及成功的真实世界迁移能力,确立了一种高效、鲁棒的隐空间主导型可扩展机器人控制范式。

原文摘要 · Abstract (English)

World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottleneck. Forcing models to reconstruct task-irrelevant visual details dissipates representational capacity and renders policies vulnerable to visual distractors. In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. Departing from the traditional reliance on visual synthesis, LeapBot-WA shifts the core of world modeling to Predictive Semantic Alignment, extracting abstract physical dynamics directly within a latent foundation space. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift. Furthermore, we design an Asymmetric Mixture-of-Transformers (MoT) architecture. During training, an Anchor Diffusion Transformer acts as a privileged dynamics expert to guide the Action Diffusion Transformer; at inference, this heavy dynamics branch is pruned, enabling zero-overhead execution. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top-tier generative WAMs on RoboTwin 2.0 without requiring large-scale trajectory pre-training. It further demonstrates superior zero-shot robustness to unseen environments and successful real-world transfer, establishing a highly efficient and robust latent-centric paradigm for scalable robotic control. Code: https://github.com/LeapWM/leapbot-wa.

机器人控制隐空间建模扩散模型零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。