用图像操作指导机器人完成厨房任务,比文字指令更快更受用户欢迎。
ImageInThat: Manipulating Images to Convey User Instructions to Robots
- 用户通过时间线式图像操作生成机器人指令
- 实验显示图像操作比文字指令快30%以上
- 适合需要直观交互的机器人任务教学场景
基础模型正显著提升机器人自主完成日常任务(如备餐)的能力,但人类指令仍不可或缺,原因包括模型性能限制、用户偏好难以捕捉以及需保留用户决策权。当前指令方式各有局限:自然语言可即时传达但易模糊抽象,而用户编程虽支持长程任务,却难准确表达意图。本文提出以图像直接操作作为新指令范式,推出名为ImageInThat的工具,让用户在时间线界面中对图像进行直观修改,自动生成机器人可执行的指令。用户研究对比了ImageInThat与文本指令方法在厨房操作任务中的表现,结果表明参与者使用ImageInThat完成任务速度更快,且更偏好该方式。补充材料(含代码)可在https://image-in-that.github.io/获取。
原文摘要 · Abstract (English)
Foundation models are rapidly improving the capability of robots in performing everyday tasks autonomously such as meal preparation, yet robots will still need to be instructed by humans due to model performance, the difficulty of capturing user preferences, and the need for user agency. Robots can be instructed using various methods-natural language conveys immediate instructions but can be abstract or ambiguous, whereas end-user programming supports longer horizon tasks but interfaces face difficulties in capturing user intent. In this work, we propose using direct manipulation of images as an alternative paradigm to instruct robots, and introduce a specific instantiation called ImageInThat which allows users to perform direct manipulation on images in a timeline-style interface to generate robot instructions. Through a user study, we demonstrate the efficacy of ImageInThat to instruct robots in kitchen manipulation tasks, comparing it to a text-based natural language instruction method. The results show that participants were faster with ImageInThat and preferred to use it over the text-based method. Supplementary material including code can be found at: https://image-in-that.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。