用视频教机器人走路,效率比现有方法快得多。
MA-ROESL: Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos
- 根据动作特征选关键帧,提升视觉语言模型生成奖励的精度。
- 三阶段训练流程使训练速度大幅提升,真实场景中也能复现行走动作。
- 适合想快速从视频学技能的机器人研究者,尤其关注效率问题。
视觉语言模型(VLM)展现出优秀的高层规划能力,可无需人工设计奖励函数,仅凭视频示范实现机器人运动技能学习。然而,现有方法存在帧采样不当和训练效率低的问题,导致计算开销大、耗时长。为此,本文提出一种基于动作感知的快速奖励优化框架(MA-ROESL),通过动作感知的帧选择机制,隐式提升VLM生成奖励函数的质量;同时采用混合三阶段训练流程,借助快速奖励优化提高训练效率,并通过在线微调获得最终策略。实验表明,MA-ROESL在仿真与真实环境中均显著提升了训练效率,且能准确复现运动技能,证明其作为高效、可扩展的视频驱动机器人运动技能学习框架的潜力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated excellent high-level planning capabilities, enabling locomotion skill learning from video demonstrations without the need for meticulous human-level reward design. However, the improper frame sampling method and low training efficiency of current methods remain a critical bottleneck, resulting in substantial computational overhead and time costs. To address this limitation, we propose Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos (MA-ROESL). MA-ROESL integrates a motion-aware frame selection method to implicitly enhance the quality of VLM-generated reward functions. It further employs a hybrid three-phase training pipeline that improves training efficiency via rapid reward optimization and derives the final policy through online fine-tuning. Experimental results demonstrate that MA-ROESL significantly enhances training efficiency while faithfully reproducing locomotion skills in both simulated and real-world settings, thereby underscoring its potential as a robust and scalable framework for efficient robot locomotion skill learning from video demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。