用无标签视频和带标签演示联合训练机器人可执行动作,提升泛化能力。
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

- 通过视觉-语言模型预测未来特征与深度变化,学习可执行的隐式动作
- 在RoboCasa上达到75.2%成功率,优于现有方法
- 适合需要少标注、强泛化的机器人控制研究者
通用机器人策略通常依赖于昂贵且难扩展的动作标注演示。相比之下,大规模人类和机器人视频虽富含物理交互信息,但常缺少可执行的机器人动作标签。本文提出WALA框架,从带动作标签的演示和无动作标签的视频中联合学习可执行的隐式动作。WALA首先在视频上预训练语义-几何隐式动作模型,通过建模当前观测与稀疏未来观测间的演化关系实现。不直接重建原始像素,而是预测DINOv3特征空间与密集深度空间中的未来增量,保留任务相关语义与几何结构,降低对外观细节的敏感性。策略训练阶段,预训练编码器提供稳定的隐式动作目标,解码器作为可训练的隐式世界模型。由机器人动作预测、隐式动作目标匹配与未来动态预测共同监督视觉-语言主干生成的隐式动作。这使得带标签演示提供可执行控制监督,而无标签视频在无需动作标注下贡献动态监督。实验表明,WALA在RoboTwin上表现优异,在RoboCasa上达到75.2%平均成功率,显著提升真实场景操作任务的策略性能与泛化能力。
原文摘要 · Abstract (English)
Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。