用视频扩散模型蒸馏知识,让机器人模型更省数据、更快推理。
Vid2WAM: Distilling Video Diffusion Priors into World Action Models

- 从大模型中提取视频先验知识,指导小模型预测未来视觉与动作。
- 在少样本专家数据下,新任务泛化能力提升,推理延迟低。
- 适合需要高效部署的机器人学习场景,尤其数据稀缺时。
世界动作模型(WAM)通过联合建模未来视觉动态和动作,提升机器人策略学习效果。然而,其可扩展性和泛化能力受限于对昂贵专家示范的依赖。本文提出Vid2WAM,一种离线蒸馏框架,将大型视频基础模型中的视觉扩散先验迁移至紧凑的WAM学生模型。给定观测和语言指令,Vid2WAM通过两条互补路径传递监督:任务条件未来轨迹直接监督学生模型的未来预测分支,而逆动力学模型则恢复具身特定的伪动作以支持动作学习。为稳健融合合成与真实监督,引入源感知残差动作适配,学习围绕共享动作主干的源特定修正,减轻噪声伪动作的干扰。推理阶段,视频教师和逆动力学模型均可移除,仅保留轻量级WAM学生模型,实现高效部署。仿真与真实世界实验表明,Vid2WAM在有限专家示范下显著提升新任务泛化能力和数据效率,同时保持低延迟推理。
原文摘要 · Abstract (English)
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by their reliance on costly expert demonstrations. We challenge this by asking whether future supervision for WAMs must originate from target-task expert trajectories. In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. Given an observation and language instruction, Vid2WAM distills supervision through two complementary channels: task-conditioned future rollouts directly supervise the student's future prediction branch, while an inverse dynamics model recovers embodiment-specific pseudo-actions for action learning. To robustly integrate synthetic and real supervision, we introduce source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from noisy pseudo-actions. During inference, both the video teacher and inverse dynamics model are discarded, leaving only the WAM student for efficient deployment. Simulation and real-world experiments demonstrate that Vid2WAM improves novel-task generalization and data efficiency under limited expert demonstrations while preserving low-latency inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。