用人类注视指导机器人精准完成复杂操作
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

- 将人眼注视转化为机器人视角的目标定位信号
- 在16个真实任务中达成最高成功率与意图准确率
- 适合需要高精度交互的智能机器人应用
视觉-语言-动作(VLA)模型虽能理解语言指令,但在实际中语言常难以精确表达意图:难以区分相似物体、确定操作位置,或应对执行中的动态变化。为此,我们提出Gaze2Act,一种利用人类注视作为动态意图信号的新框架。该方法通过跨视角语义匹配,将第一人称注视映射至机器人视角,生成物体掩码和注视点,实现从粗到精的目标指定。这些线索通过感知层提示与动作层条件化融入策略,使机器人能聚焦相关区域并执行精准操作。在七类任务、16个真实机器人任务(基于Unitree G1人形机器人)上系统评估显示,Gaze2Act在意图准确率与任务成功率上均达到当前最优。尤其在物体辨识、细粒度交互与动态意图调控方面显著优于基线。结果表明,人类注视是一种自然、低负担且高度表达性的闭环控制模态。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level conditioning, allowing the robot to attend to relevant regions and execute precise interactions under dynamic intent. In a systematic evaluation across seven task categories and 16 real-robot tasks on a Unitree G1 humanoid, Gaze2Act achieves state-of-the-art performance in both intent accuracy and task success rate. It notably outperforms baselines in object disambiguation, fine-grained interaction, and dynamic intent steering. These results demonstrate that human gaze provides a natural, low-burden, and highly expressive modality for human-in-the-loop VLA control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。