用隐式动作混合让机器人从想象未来视频直接执行动作
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

- 用预训练逆动力学模型提取生成视频中的隐式动作
- 在仿真和真实机器人上提升成功率与动作一致性
- 适合做视觉想象与动作执行衔接的研究者
视频生成模型能预测长时序未来视觉状态,为机器人操作提供想象能力,但如何将这些想象用于实际动作执行仍具挑战。现有方法或基于预测帧条件化策略,或直接将视频解码为动作,均存在视觉真实度与控制相关性不匹配的问题。这导致预测结果偏向感知保真,而非驱动状态转移的动作本质,造成控制间接且不稳定。为此,我们提出MoLA(Mixture of Latent Actions),一种面向控制的接口,将想象的未来视频转化为可执行的表示。不同于直接传递预测帧,MoLA利用一组预训练逆动力学模型,推断生成视觉变化所隐含的多模态隐式动作。这些模态感知的逆动力学模型融合语义、深度和光流信息,提供结构化且物理合理的动作表征,有效连接视频想象与策略执行。我们在模拟基准(LIBERO、CALVIN、LIBERO-Plus)和真实机器人任务中验证该方法,显著提升任务成功率、时间一致性和泛化能力。
原文摘要 · Abstract (English)
Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or directly decode generated videos into actions, both suffering from a mismatch between visual realism and control relevance. As a result, predicted observations emphasize perceptual fidelity rather than action-centric causes of state transitions, leading to indirect and unstable control. To address this gap, we propose MoLA (Mixture of Latent Actions), a control-oriented interface that transforms imagined future videos into executable representations. Instead of passing predicted frames directly to the policy, MoLA leverages a mixture of pretrained inverse dynamics models to infer a mixture of latent actions implied by generated visual transitions. These modality-aware inverse dynamics models capture complementary semantic, depth, and flow cues, providing a structured and physically grounded action representation that bridges video imagination and policy execution. We evaluate our approach on simulated benchmarks (LIBERO, CALVIN, and LIBERO-Plus) and real-world robot manipulation tasks, achieving consistent gains in task success, temporal consistency, and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。