无需预先学习意图,仅靠视觉就能实时理解用户目标并协助完成操作。
Toward Zero-Shot User Intent Recognition in Shared Autonomy
- 基于末端视觉的零样本意图识别,不依赖历史数据或演示。
- 在未知物体位置下,性能接近拥有先验知识的最优模型。
- 适合人机协作场景,尤其适用于意图动态变化的任务。
共享自主的核心挑战在于:如何让高自由度机器人通过推断用户意图来辅助而非干扰人类完成任务。现有方法通常需要事先掌握所有可能的人类意图,或大量与用户交互才能学习。本文提出一种零样本、纯视觉的共享自主框架(VOSA),仅依赖末端视觉,即可在未知且动态变化的物体位置下实时估计用户意图,并结合混合控制策略帮助用户完成操作任务。我们在Kinova Gen3机械臂上实现简化版VOSA,在三个桌面操作任务中开展用户研究。结果表明,VOSA性能接近具备先验知识的最优基线模型,同时显著优于无人辅助的遥操作。在真实场景中,当可选意图完全或部分未知时,VOSA所需的人力和时间更少,且多数参与者更偏好该系统。实验验证了使用现成视觉算法实现灵活高效人机协作的可行性。代码与视频见:https://sites.google.com/view/zeroshot-sharedautonomy/home。
原文摘要 · Abstract (English)
A fundamental challenge of shared autonomy is to use high-DoF robots to assist, rather than hinder, humans by first inferring user intent and then empowering the user to achieve their intent. Although successful, prior methods either rely heavily on a priori knowledge of all possible human intents or require many demonstrations and interactions with the human to learn these intents before being able to assist the user. We propose and study a zero-shot, vision-only shared autonomy (VOSA) framework designed to allow robots to use end-effector vision to estimate zero-shot human intents in conjunction with blended control to help humans accomplish manipulation tasks with unknown and dynamically changing object locations. To demonstrate the effectiveness of our VOSA framework, we instantiate a simple version of VOSA on a Kinova Gen3 manipulator and evaluate our system by conducting a user study on three tabletop manipulation tasks. The performance of VOSA matches that of an oracle baseline model that receives privileged knowledge of possible human intents while also requiring significantly less effort than unassisted teleoperation. In more realistic settings, where the set of possible human intents is fully or partially unknown, we demonstrate that VOSA requires less human effort and time than baseline approaches while being preferred by a majority of the participants. Our results demonstrate the efficacy and efficiency of using off-the-shelf vision algorithms to enable flexible and beneficial shared control of a robot manipulator. Code and videos available here: https://sites.google.com/view/zeroshot-sharedautonomy/home.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。