用眼神+大模型识别用户隐含意图,让助手机器人更懂你。
MindEye-OmniAssist: A Gaze-Driven LLM-Enhanced Assistive Robot System for Implicit Intention Recognition and Task Execution
- 通过眼神输入结合大模型和视觉基础模型,实现开放词汇的意图识别。
- 在未定义任务中达到41/55的成功率,支持复杂多样化操作。
- 适合需要自然交互的残障人士辅助场景,提升系统智能与适应性。
在辅助机器人系统中,基于视线的交互具有广阔前景。然而,现有系统主要支持基础抓取动作,功能受限;且意图识别能力不足,难以提供多样化帮助。本文提出一种由大语言模型(LLM)和视觉基础模型(VFM)驱动的开放隐含意图识别框架,可处理视线输入并识别非预设场景下的用户意图。进一步,我们构建了基于视线的增强型助手机器人系统(MindEye-OmniAssist),通过开放词汇目标检测器、意图识别网络与大语言模型,推断用户完整意图,并结合眼动反馈生成动作序列完成任务。真实世界实验表明,该系统在多种未定义任务中取得了41/55的整体成功率。初步结果证明,该方法具备显著提升辅助系统通用性与有效性潜力,为用户提供更友好的人机交互界面。
原文摘要 · Abstract (English)
A promising effective human-robot interaction in assistive robotic systems is gaze-based control. However, current gaze-based assistive systems mainly help users with basic grasping actions, offering limited support. Moreover, the restricted intent recognition capability constrains the assistive system's ability to provide diverse assistance functions. In this paper, we propose an open implicit intention recognition framework powered by Large Language Model (LLM) and Vision Foundation Model (VFM), which can process gaze input and recognize user intents that are not confined to predefined or specific scenarios. Furthermore, we implement a gaze-driven LLM-enhanced assistive robot system (MindEye-OmniAssist) that recognizes user's intentions through gaze and assists in completing task. To achieve this, the system utilizes open vocabulary object detector, intention recognition network and LLM to infer their full intentions. By integrating eye movement feedback and LLM, it generates action sequences to assist the user in completing tasks. Real-world experiments have been conducted for assistive tasks, and the system achieved an overall success rate of 41/55 across various undefined tasks. Preliminary results show that the proposed method holds the potential to provide a more user-friendly human-computer interaction interface and significantly enhance the versatility and effectiveness of assistive systems by supporting more complex and diverse task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。