用可解释的动作图像实现机器人端到端控制,零样本表现更优。
Action Images: End-to-End Policy Learning via Multiview Video Generation

- 将7自由度动作转为多视角像素级动作视频,直接输入视频模型
- 在RLBench上零样本成功率最高,视频-动作联合生成质量更优
- 适合做视觉驱动的机器人控制与跨场景迁移研究
世界动作模型(WAMs)为机器人策略学习提供了新方向,能利用强大的视频骨干网络建模未来状态。但现有方法常依赖独立的动作模块,或使用非像素对齐的动作表示,难以充分调用预训练视频模型的知识,限制了跨视角和环境的迁移能力。本文提出统一的世界动作模型——动作图像(Action Images),将策略学习转化为多视角视频生成任务。不同于低维令牌编码,我们将7自由度机器人动作转化为可解释的动作图像:以2D像素为基准的多视角动作视频,显式追踪机械臂运动。这种像素对齐的动作表示使视频骨干网络本身即可作为零样本策略,无需额外策略头或动作模块。除控制外,同一模型还支持视频-动作联合生成、动作条件视频生成及动作标注,共享统一表示。在RLBench和真实世界评估中,该模型实现了最强的零样本成功率,并显著提升视频-动作联合生成质量,表明可解释的动作图像是策略学习的有力路径。
原文摘要 · Abstract (English)
World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate action modules, or use action representations that are not pixel-grounded, making it difficult to fully exploit the pretrained knowledge of video models and limiting transfer across viewpoints and environments. In this work, we present Action Images, a unified world action model that formulates policy learning as multiview video generation. Instead of encoding control as low-dimensional tokens, we translate 7-DoF robot actions into interpretable action images: multi-view action videos that are grounded in 2D pixels and explicitly track robot-arm motion. This pixel-grounded action representation allows the video backbone itself to act as a zero-shot policy, without a separate policy head or action module. Beyond control, the same unified model supports video-action joint generation, action-conditioned video generation, and action labeling under a shared representation. On RLBench and real-world evaluations, our model achieves the strongest zero-shot success rates and improves video-action joint generation quality over prior video-space world models, suggesting that interpretable action images are a promising route to policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。