用视觉语言模型看人类演示视频,自动生成机器人操作计划。
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
- 通过关键帧选取+视觉感知+VLM推理构建端到端解析流水线。
- 在三类长时序抓取放置任务上超越现有视频输入VLM基线。
- 可直接部署于仿真环境与真实机械臂,具备实用价值。
视觉语言模型(VLM)因其常识推理和泛化能力被引入机器人领域,用于从自然语言指令生成任务与运动规划,并模拟训练数据。本文探索利用VLM解析人类示范视频并生成机器人任务规划。提出SeeDo框架,将关键帧选择、视觉感知与VLM推理集成至统一流程中,使VLM能“看”懂人类示范并生成机器人可执行的计划。为验证方法有效性,收集了三类多样化场景下的长时序人类示范视频,设计多维度评估指标,对比包括先进视频输入VLM在内的多个基线。实验表明SeeDo性能更优。进一步在仿真环境与真实机械臂上部署生成的任务计划,验证了其可行性与实用性。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have recently been adopted in robotics for their capability in common sense reasoning and generalizability. Existing work has applied VLMs to generate task and motion planning from natural language instructions and simulate training data for robot learning. In this work, we explore using VLM to interpret human demonstration videos and generate robot task planning. Our method integrates keyframe selection, visual perception, and VLM reasoning into a pipeline. We named it SeeDo because it enables the VLM to ''see'' human demonstrations and explain the corresponding plans to the robot for it to ''do''. To validate our approach, we collected a set of long-horizon human videos demonstrating pick-and-place tasks in three diverse categories and designed a set of metrics to comprehensively benchmark SeeDo against several baselines, including state-of-the-art video-input VLMs. The experiments demonstrate SeeDo's superior performance. We further deployed the generated task plans in both a simulation environment and on a real robot arm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。