用眼神+大模型让机器人自动理解用户意图并执行操作
Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models
- 通过眼动追踪与视觉语言模型结合,将视线聚焦转化为语义意图
- 在桌面上多种任务中实现无需特定训练的自主操作,成功率显著提升
- 适合助老助残场景,为自然交互提供可扩展的智能控制方案
设计直观的机器人控制界面仍是实现高效人机交互的核心挑战,尤其在辅助护理场景中。眼动提供快速、非侵入且富含意图的信息,是传达用户目标的理想方式。本文提出GAMMA(Gaze Assisted Manipulation for Modular Autonomy),利用自我中心眼动追踪和视觉-语言模型,推断用户意图并自主执行机器人操作任务。通过将眼动注视点置于场景上下文中,系统将视觉注意力映射为高层次语义理解,实现技能选择与参数化,无需针对具体任务进行训练。我们在一系列桌面操作任务上评估GAMMA,对比了无推理能力的基线眼动控制方法。结果表明,GAMMA提供了稳健、直观且泛化的控制能力,凸显了结合基础模型与眼动输入在实现自然、可扩展机器人自主性方面的潜力。
原文摘要 · Abstract (English)
Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality, making it an attractive channel for conveying user goals. In this work, we present GAMMA (Gaze Assisted Manipulation for Modular Autonomy), a system that leverages ego-centric gaze tracking and a vision-language model to infer user intent and autonomously execute robotic manipulation tasks. By contextualizing gaze fixations within the scene, the system maps visual attention to high-level semantic understanding, enabling skill selection and parameterization without task-specific training. We evaluate GAMMA on a range of table-top manipulation tasks and compare it against baseline gaze-based control without reasoning. Results demonstrate that GAMMA provides robust, intuitive, and generalizable control, highlighting the potential of combining foundation models and gaze for natural and scalable robot autonomy. Project website: https://gamma0.vercel.app/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。