用无动作视频学机器人运动先验,提升少样本学习效果
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
- 将关键点轨迹转为离散运动编码,分离视觉预测与动作推断
- 在低数据下性能提升1.2-2.2倍,像素预测准确率超前代2.5倍
- 适合从人类视频中学习控制策略的科研与工程人员
机器人任务中带动作标注的数据稀缺且昂贵,限制了策略泛化能力。相比之下,大量无动作标注的视频数据易获取,但如何从中提取有效策略仍是挑战。我们提出AMPLIFY框架,通过关键点轨迹生成紧凑的离散运动令牌,将视觉动态编码为可复用的运动先验。该模块化设计将视觉运动预测与动作推理解耦,使两者可独立扩展:在海量无动作视频上训练前向动力学模型,在少量标注数据上训练逆动力学模型。实验表明,所学动态模型精度显著提升,最大降低3.7倍均方误差,像素预测准确率提升超2.5倍;在下游策略学习中,低数据场景下性能提升1.2-2.2倍,仅用人类无动作视频即实现1.4倍平均增益,并首次实现从零分布动作数据泛化到LIBERO任务。此外,该模型还可作为通用潜空间世界模型,增强视频预测质量。结果展示了利用异构数据构建高效、泛化性强世界模型的新范式。
原文摘要 · Abstract (English)
Action-labeled data for robotics is scarce and expensive, limiting the generalization of learned policies. In contrast, vast amounts of action-free video data are readily available, but translating these observations into effective policies remains a challenge. We introduce AMPLIFY, a novel framework that leverages large-scale video data by encoding visual dynamics into compact, discrete motion tokens derived from keypoint trajectories. Our modular approach separates visual motion prediction from action inference, decoupling the challenges of learning what motion defines a task from how robots can perform it. We train a forward dynamics model on abundant action-free videos and an inverse dynamics model on a limited set of action-labeled examples, allowing for independent scaling. Extensive evaluations demonstrate that the learned dynamics are both accurate, achieving up to 3.7x better MSE and over 2.5x better pixel prediction accuracy compared to prior approaches, and broadly useful. In downstream policy learning, our dynamics predictions enable a 1.2-2.2x improvement in low-data regimes, a 1.4x average improvement by learning from action-free human videos, and the first generalization to LIBERO tasks from zero in-distribution action data. Beyond robotic control, we find the dynamics learned by AMPLIFY to be a versatile latent world model, enhancing video prediction quality. Our results present a novel paradigm leveraging heterogeneous data sources to build efficient, generalizable world models. More information can be found at https://amplify-robotics.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。