通过学习通用潜动作提升机器人操控的少样本迁移能力
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- 联合预测未来帧和动作轨迹,融入物理先验知识
- 仅需每任务10条真实数据即完成5个挑战任务
- 适合追求少样本迁移的机器人学习研究者
从大规模物体操作视频中学习可迁移的潜动作表征,能显著提升下游机器人任务的泛化能力,因为这些表征与具体机器人形态无关。现有方法主要依赖视觉重建目标,忽视物理先验,导致通用表征学习效果不佳。为此,我们提出一种通用潜动作学习框架,以任务指令和多帧图像为输入,同时优化未来帧重建和动作序列预测。与以往工作不同,引入动作预测(如夹爪或手部轨迹与朝向)使模型能捕捉真实世界中的距离、方向等物理先验,从而实现无缝迁移。我们进一步将潜动作分解为可学习的运动标记和场景标记,以区分机器人主动运动与环境变化,过滤无关动态。通过将学习到的潜动作蒸馏至最新VLA模型,我们在模拟(SIMPLER和LIBERO)和真实机器人场景中均取得优异表现。值得注意的是,仅需在Franka机器人上收集每任务10条真实轨迹,本方法即可成功完成全部五个挑战任务,展现出强大的少样本迁移能力。
原文摘要 · Abstract (English)
Learning transferable latent actions from large-scale object manipulation videos can significantly enhance generalization in downstream robotics tasks, as such representations are agnostic to different robot embodiments. Existing approaches primarily rely on visual reconstruction objectives while neglecting physical priors, leading to sub-optimal performance in learning universal representations. To address these challenges, we propose a Universal Latent Action Learning framework that takes task instructions and multiple frames as inputs, and optimizes both future frame reconstruction and action sequence prediction. Unlike prior works, incorporating action predictions (e.g., gripper or hand trajectories and orientations) allows the model to capture richer physical priors such as real-world distances and orientations, thereby enabling seamless transferability to downstream tasks. We further decompose the latent actions into learnable motion and scene tokens to distinguish the robot's active movements from environmental changes, thus filtering out irrelevant dynamics. By distilling the learned latent actions into the latest VLA models, we achieve strong performance across both simulated (SIMPLER and LIBERO) and real-world robot settings. Notably, with only 10 real-world trajectories per task collected on a Franka robot, our approach successfully completes all five challenging tasks, demonstrating strong few-shot transferability in robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。