arXiv:2608.03701cs.ROcs.AI2026-08被引 2

轻量级模型实现机器人任务的端到端未来状态预测与动作生成。

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

论文配图:LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
图 1 · 摘自论文原文
  • 在紧凑隐空间中联合预测未来状态与生成动作,无需多阶段训练。
  • 单卡24GB GPU训练下,在50个RoboTwin任务上达成90.48%成功率。
  • 提出无需语言的视觉过渡令牌,以方向编码任务意图,适合真实机器人部署。

世界-动作建模(WAM)已成为机器人控制的前沿范式,使模型能超越对观测的被动响应,预判场景演化。然而现有方法常伴随高昂计算开销:像素空间方法过度关注无关视觉细节,部分隐空间方法需多阶段训练构建推理空间,导致训练成本高,难以在有限算力下应用。本文提出轻量级隐空间推理世界-动作模型LiLa-WAM,可在单张24GB GPU上端到端训练,其核心是融合未来状态预测与动作生成的紧凑隐空间,兼顾轻量化与控制对齐性。为任务表征,我们设计视觉过渡令牌(VTT),将任务编码为视觉特征空间中的方向,无需语言输入。在RoboTwin 2.0、LIBERO及真实机器人任务上的实验表明,该模型在50个RoboTwin任务中取得90.48%的成功率,验证了其高效性与实用性。

原文摘要 · Abstract (English)

World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.

机器人控制隐空间建模端到端训练轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。