arXiv:2606.15768cs.ROcs.AI2026-06被引 21

用压缩的潜在视觉目标提升机器人动作预测效率

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

论文配图:LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
图 1 · 摘自论文原文
  • 用预训练视觉模型的潜在空间生成未来场景特征,替代高耗时像素级视频生成
  • 在LIBERO数据集上达到98.6%成功率,推理速度比像素空间模型快24倍
  • 适合需要快速响应的机器人控制任务,尤其适用于真实世界操作场景

视觉-语言-动作模型(VLAs)利用大规模视觉-语言预训练实现语义机器人控制,但通常缺乏对动作如何改变环境的显式预见能力。世界-动作模型(WAMs)通过基于预测未来来弥补这一不足,但现有方法多依赖计算成本高昂的视频生成,存在大量像素级冗余。本文提出LaWAM,一种潜在世界动作模型,通过紧凑的潜在视觉子目标将预测动态暴露给机器人策略,而非重建未来视频。核心是潜在动作条件下的潜在世界模型(LaWM),通过在预训练视觉基础模型的潜在空间中训练潜在动作模型,并复用其前向解码器预测场景演化的未来观测特征。LaWAM随后以这些预测的潜在视觉子目标为条件生成动作,实现动态感知的机器人控制。在LIBERO(98.6%成功率)、RoboTwin(91.22%成功率)和真实世界操作任务中均达到当前最优或竞争力表现,同时保持低延迟推理。每动作块预测仅需187毫秒,相比像素空间WAMs,墙钟延迟最高降低24倍。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.

机器人控制潜在建模高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。