用手绘草图直接控制移动操作机器人,无需额外设备。
Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches
- 用视觉语言模型解析叠加在画面中的手绘草图,理解其意图。
- 可精准识别7类任务、5种草图形状,支持抓取角度等细节控制。
- 用户实验表明,相比现有2D界面,操作更直观,易用性更强。
为使助手机器人融入日常生活,使用如二维设备等常见工具实现远程操控至关重要。手绘草图是通过二维设备直观控制机器人的有效方式。然而,由于相似草图在不同场景中含义不同,现有方法需依赖额外模态来确定语义,操作复杂且降低可用性。本文提出Sketch-MoMa,一种基于用户手绘草图指令的移动操作机器人遥操作系统。利用视觉语言模型(VLMs)理解叠加在观测图像上的草图,推断出绘制形状与机器人低级任务。结合草图与生成形状,完成识别与运动规划,实现精确、直观的操作。我们在7项任务和5种草图形状上验证了该方法的有效性,并证明其能准确指定具体动作,如抓取方式与旋转角度。用户实验(14名参与者)显示,该方法在可用性上具有竞争力。
原文摘要 · Abstract (English)
To use assistive robots in everyday life, a remote control system with common devices, such as 2D devices, is helpful to control the robots anytime and anywhere as intended. Hand-drawn sketches are one of the intuitive ways to control robots with 2D devices. However, since similar sketches have different intentions from scene to scene, existing work needs additional modalities to set the sketches' semantics. This requires complex operations for users and leads to decreasing usability. In this paper, we propose Sketch-MoMa, a teleoperation system using the user-given hand-drawn sketches as instructions to control a robot. We use Vision-Language Models (VLMs) to understand the user-given sketches superimposed on an observation image and infer drawn shapes and low-level tasks of the robot. We utilize the sketches and the generated shapes for recognition and motion planning of the generated low-level tasks for precise and intuitive operations. We validate our approach using state-of-the-art VLMs with 7 tasks and 5 sketch shapes. We also demonstrate that our approach effectively specifies the detailed motions, such as how to grasp and how much to rotate. Moreover, we show the competitive usability of our approach compared with the existing 2D interface through a user experiment with 14 participants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。