用户画图教机器人任务,比动手操作更省力高效。
L2D2: Robot Learning from 2D Drawings
- 用平板在2D图像上画轨迹,替代物理引导机器人
- 结合视觉语言分割自动重置场景,生成多样演示数据
- 少量真实示范+大量画图,提升策略性能与长程泛化
机器人应从人类获取新任务。但人类如何传达意图?现有方法依赖物理引导,随数据量增加,人工操作和环境重置成本过高。本文提出L2D2:通过绘图界面与模仿学习算法,用户在2D图像上绘制并标注轨迹即可示范任务。系统利用视觉-语言分割技术自动变换物体位置,生成合成图像供用户绘图,无需重复物理重置。为弥补2D绘图信息不足的问题,L2D2引入少量真实物理示范,将静态2D绘图映射到动态3D世界。实验与用户研究显示,相比传统方式,L2D2显著降低时间与精力消耗,用户更偏好绘图;相较于其他绘图方法,其训练所需数据更少、策略表现更好,并能推广至更长时序任务。
原文摘要 · Abstract (English)
Robots should learn new tasks from humans. But how do humans convey what they want the robot to do? Existing methods largely rely on humans physically guiding the robot arm throughout their intended task. Unfortunately -- as we scale up the amount of data -- physical guidance becomes prohibitively burdensome. Not only do humans need to operate robot hardware but also modify the environment (e.g., moving and resetting objects) to provide multiple task examples. In this work we propose L2D2, a sketching interface and imitation learning algorithm where humans can provide demonstrations by drawing the task. L2D2 starts with a single image of the robot arm and its workspace. Using a tablet, users draw and label trajectories on this image to illustrate how the robot should act. To collect new and diverse demonstrations, we no longer need the human to physically reset the workspace; instead, L2D2 leverages vision-language segmentation to autonomously vary object locations and generate synthetic images for the human to draw upon. We recognize that drawing trajectories is not as information-rich as physically demonstrating the task. Drawings are 2-dimensional and do not capture how the robot's actions affect its environment. To address these fundamental challenges the next stage of L2D2 grounds the human's static, 2D drawings in our dynamic, 3D world by leveraging a small set of physical demonstrations. Our experiments and user study suggest that L2D2 enables humans to provide more demonstrations with less time and effort than traditional approaches, and users prefer drawings over physical manipulation. When compared to other drawing-based approaches, we find that L2D2 learns more performant robot policies, requires a smaller dataset, and can generalize to longer-horizon tasks. See our project website: https://collab.me.vt.edu/L2D2/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。