用在线蒸馏修复视频优先世界模型的行动能力缺陷
WAM-OPD: On-Policy Distillation for World Action Models

- 学生在环境中行动生成历史,教师用一致视频与动作目标监督
- 在手部交接和放物品抽屉任务中成功率分别提升至58.3%和33.3%
- 无需稀疏奖励强化学习,适合部署后微调视频优先世界模型
世界动作模型(WAM)将视觉未来预测与机器人动作生成结合,但加速的学生模型在蒸馏过程中可能丢失任务能力,且后续遇到离线数据覆盖不足的状态。本文研究在线蒸馏(OPD)能否在不依赖稀疏奖励强化学习的前提下修复此类问题。提出WAM-OPD,一种适用于视频优先型WAM的部署一致后训练方法。学生在环境中的行动决定历史分布,冻结的教师为这些学生生成的历史提供连贯的视频与动作目标标签;学生动作分支在自身生成的视频计划下进行训练,与实际部署一致。联合视频与动作损失更新共享主干中的轻量适配器,同时引入动作流匹配正则化。在RoboTwin 2.0初步实验中,单视频/单动作步的Flash-WAM在手部交接任务上成功率从0.0%提升至58.3%,在放物品抽屉任务上从16.7%提升至33.3%。这些任务特定结果为能力验证,非广泛泛化证据,但仍表明对学生生成历史进行密集教师监督是视频优先型WAM有前景的后训练接口。
原文摘要 · Abstract (English)
World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. We study whether on-policy distillation (OPD) can repair such a student without requiring sparse-reward reinforcement learning. We introduce WAM-OPD, a deployment-consistent post-training recipe for a video-first WAM. The student acts in the environment and therefore determines the history distribution. A frozen teacher labels those student histories with coherent video and action targets, while the student action branch is trained under its own generated video plan, as it is at deployment. Joint video and action losses update lightweight adapters in the shared backbone, together with an action flow-matching regularizer. In preliminary RoboTwin 2.0 studies on two tasks, the released one-video/one-action-step Flash-WAM improves from 0.0% to 58.3% success on HANDOVER MIC, and from 16.7% to 33.3% on PUT OBJECT CABINET. These task-specific results are an initial capability proof rather than evidence of broad or uniform generalization. They nevertheless suggest that dense teacher supervision on student-induced histories is a promising post-training interface for video-first WAMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。