仅用一段真人视频,让机器人学会重复性摆放任务
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
- 用视觉基础模型+新型槽位检测器,从单段视频中学习任务
- 在真实视频上测试,摆放准确率显著优于基线方法
- 适合需要快速教新任务的工业自动化场景
当前多数机器人学习方法聚焦于预定义任务,难以泛化到新任务,扩展技能需大量训练数据。本文针对重复性任务(如打包)的机器人教学问题,提出仅需一段真人操作视频即可实现新任务学习。系统需理解视频中被拾取的物体(抓取物)及目标放置槽位,并在推理时重新识别两者及其相对位姿以执行动作。为此,我们提出SLeRP模块化系统,融合多个先进视觉基础模型与新型槽位级放置检测器Slot-Net,无需昂贵视频数据训练。我们在新构建的真实世界视频基准上评估系统,结果表明SLeRP性能优于多个基线方法,且可部署于真实机器人。
原文摘要 · Abstract (English)
The majority of modern robot learning methods focus on learning a set of pre-defined tasks with limited or no generalization to new tasks. Extending the robot skillset to novel tasks involves gathering an extensive amount of training data for additional tasks. In this paper, we address the problem of teaching new tasks to robots using human demonstration videos for repetitive tasks (e.g., packing). This task requires understanding the human video to identify which object is being manipulated (the pick object) and where it is being placed (the placement slot). In addition, it needs to re-identify the pick object and the placement slots during inference along with the relative poses to enable robot execution of the task. To tackle this, we propose SLeRP, a modular system that leverages several advanced visual foundation models and a novel slot-level placement detector Slot-Net, eliminating the need for expensive video demonstrations for training. We evaluate our system using a new benchmark of real-world videos. The evaluation results show that SLeRP outperforms several baselines and can be deployed on a real robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。