让机器人在行动前先想象后果,提升决策准确性。
$ω$-EVA: Envision, Verify, and Act with Latent Interactive World Models

- 构建可交互的潜在世界模型,实现想象-验证-执行闭环。
- 在多种仿真环境中,动作成功率显著优于基线方法。
- 无需额外机器人数据预训练,适合实际部署应用。
具身策略通常将当前观测直接映射为动作,隐含了动作后果。世界模型虽能提供预测监督或外部模拟,但很少允许策略在行动前审视自身提议的后果。我们提出 $ω$-EVA,一种潜在交互式世界模型,实现具身动作生成的“想象—验证—执行”循环。其三阶段框架学习动作条件下的潜在动态,基于动态感知的视觉表征训练语言条件流策略,并将策略提案反馈至世界模型。三分支精炼器联合推理当前状态、条件未来与提议动作,生成最终动作块。由于后果推理保持在潜在特征空间,$ω$-EVA 在推理时不生成未来视频。在单臂、双臂、长时程及扰动等多种仿真设置中评估表明,完整交互流程持续提升提议策略性能,潜在诊断显示有意义的动作条件未来结构。参数量约1.2B,无需额外机器人数据预训练,展现出紧凑且竞争力强的性能-规模-数据权衡,使世界模型成为主动的动作反馈模块而非被动预测器。
原文摘要 · Abstract (English)
Embodied policies typically map current observations directly to actions, leaving candidate-action consequences implicit. World models provide predictive supervision, representations, or external simulation, but rarely let a policy inspect the imagined consequence of its own proposal before acting. We introduce $ω$-EVA, a latent interactive world model that realizes an Envision--Verify--Act loop for embodied action generation. Its three-stage framework learns action-conditioned latent dynamics, trains a language-conditioned flow policy on dynamics-aware visual representations, and feeds the policy's proposal back through the world model. A tri-branch refiner jointly reasons over the current state, proposal-conditioned future, and proposed action to produce the final action chunk. Because consequence reasoning remains in latent feature space, $ω$-EVA avoids generating future videos at inference. Evaluations across diverse single-arm, bimanual, long-horizon, and perturbed simulation settings show that the complete interaction pipeline consistently improves the proposal policy, while latent diagnostics indicate meaningful action-conditioned future structure. With approximately 1.2B parameters and no additional robot-data pretraining, $ω$-EVA demonstrates a compact and competitive performance--scale--data trade-off, making the world model an active action-feedback module rather than a passive predictor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。