用少量人类视频让机器人学会精细动作,效果远超现有方法
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- 用基础模型直接从人类视频学细粒度动作,无需预设动作模板
- 单视频即达强性能,五视频时比其他方法提升超39%以上
- 特别适合高精度和长序列任务,真实场景表现显著更好
视觉模仿学习(VIL)为机器人系统高效获取新技能提供了直观策略。近年来,基础模型特别是视觉语言模型(VLMs)在视觉与语言推理方面展现出强大能力,适用于VIL任务。然而,现有方法主要利用这些模型从人类示范中学习高层计划,仍依赖预定义的动作基元来执行物理交互,成为机器人系统的瓶颈。本文提出FMimic,一种新范式,仅需少量人类视频即可让基础模型直接学习可泛化的细粒度动作技能。大量实验表明,FMimic在仅使用一个视频时即表现出色,使用五个视频时显著优于所有其他方法。在RLBench多任务实验和真实世界操作任务中,分别实现超过39%和29%的性能提升;在高精度任务上超越基线34%以上,在长时序任务中提升达47%。
原文摘要 · Abstract (English)
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in foundation models, particularly Vision Language Models (VLMs), have demonstrated remarkable capabilities in visual and linguistic reasoning for VIL tasks. Despite this progress, existing approaches primarily utilize these models for learning high-level plans from human demonstrations, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck for robotic systems. In this work, we present FMimic, a novel paradigm that harnesses foundation models to directly learn generalizable skills at even fine-grained action levels, using only a limited number of human videos. Extensive experiments demonstrate that our FMimic delivers strong performance with a single human video, and significantly outperforms all other methods with five videos. Furthermore, our method exhibits significant improvements of over 39% and 29% in RLBench multi-task experiments and real-world manipulation tasks, respectively, and exceeds baselines by more than 34% in high-precision tasks and 47% in long-horizon tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。