arXiv:2608.22067cs.ROcs.AI2026-08

不生成视频,直接从未来状态推断机器人动作,更高效准确。

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

论文配图:DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
图 1 · 摘自论文原文
  • 跳过视频生成,直接从未来隐状态推动作,减少冗余计算。
  • 在4个长时序操作任务上实现62.5%的任务成功率和81.3%的进度率。
  • 适合追求低延迟、高效率的真实机器人控制场景。

世界-动作模型(WAMs)基于视频生成骨干网络构建机器人控制,联合预测密集的未来视觉轨迹和机器人动作。我们认为,视频生成是世界-动作建模中不必要的中间目标。对于机器人操作,世界模型的目标并非复现每个中间时刻的世界外观,而是预测动作执行后世界将达到的状态。中间帧仅描述物理状态间的视觉过渡,消耗大量模型容量与计算资源,却无法直接指明机器人动作所期望的物理结果。本文提出DELE-w0.5,通过预测的未来状态直接推断机器人动作,无需依赖视频生成。具体而言,DELE-w0.5从对应的紧凑未来隐状态中推断动作序列。该未来隐状态捕捉了动作相关的物理结果,并作为世界建模与动作生成之间的显式桥梁。DELE-w0.5的核心设计原则是建模机器人动作下物理世界的变化,而非视觉外观逐帧演化。这一设定消除了密集视频表示引入的高维视觉冗余,从而实现更低的训练成本和更低的推理延迟。在四个长时序操作任务上的640次真实机器人实验中,DELE-w0.5在所有对比策略中表现最佳,总体任务成功率达62.5%,宏有序阶段进展达81.3%。其任务成功率比最强基线高出32.5个百分点,宏进展领先20.1个百分点。

原文摘要 · Abstract (English)

World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 640 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress. It outperforms the strongest baseline by 32.5 percentage points in full-task success and 20.1 percentage points in macro progress.

机器人控制动作推断隐状态建模高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。