用少量人类视频让视觉语言模型直接学习精细动作,提升机器人模仿能力。
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
- 基于人类视频提取物体中心运动,通过分层约束学习精细动作
- 仅用5个视频,在真实任务中提升27%以上表现
- 适合需要少样本学习的机器人精细化操作场景
视觉模仿学习(VIL)为机器人系统获取新技能提供了一种高效且直观的策略。近期视觉语言模型(VLMs)在视觉与语言推理方面展现出卓越性能,适用于VIL任务。然而,现有方法仍简单地将VLM用于从人类视频中学习高层规划,并依赖预定义的动作基元执行物理交互,这成为主要瓶颈。本文提出VLMimic,一种新范式,使VLM能够仅凭少量人类视频直接学习精细动作层级。具体而言,VLMimic首先从人类视频中定位物体中心运动,利用分层约束表示学习技能,从而在有限人类视频基础上推导出精细动作。这些技能通过迭代对比策略进行优化与更新,实现对未见环境的高效适应。大量实验表明,仅使用5个视频,VLMimic在RLBench和真实世界操作任务中分别提升超过27%和21%,在长时序任务中超越基线37%以上。
原文摘要 · Abstract (English)
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by over 37% in long-horizon tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。