用人类注视行为学习意图,让机器人更懂人想做什么。
GazeVLA: Learning Human Intention for Robotic Manipulation

- 以注视作为意图代理,预训练+微调框架
- 少样本下仍表现优异,长任务成功率超基线
- 适合需要理解人类意图的协作机器人场景
具身基础模型在机器人操作中取得显著进展,但仍严重依赖大规模机器人示范数据。尽管近期研究尝试利用人类数据缓解这一依赖,但因人与机器人之间的固有具身差距,有效提取可迁移知识仍面临挑战。本文认为,人类动作背后的意图可作为桥接该差距的有力中间表征。我们提出一种新框架,显式学习并传递人类意图以促进机器人操作。具体地,通过注视建模意图,因其自然先于物理动作,且可观测。模型首先在大规模第一人称人类数据集上预训练,捕捉意图与动作的协同关系,随后在少量机器人和人类数据上微调。推理时采用思维链范式,先预测意图再执行动作。在仿真和真实世界环境中,针对长周期、细粒度任务,以及少样本和鲁棒性基准的广泛评估显示,本方法持续优于强基线,泛化能力更强,达到当前最优性能。
原文摘要 · Abstract (English)
Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works have explored leveraging human data to alleviate this dependency, effectively extracting transferable knowledge remains a significant challenge due to the inherent embodiment gap between human and robot. We argue that the intention underlying human actions can serve as a powerful intermediate representation for bridging this gap. In this paper, we introduce a novel framework that explicitly learns and transfers human intention to facilitate robotic manipulation. Specifically, we model intention through gaze, as it naturally precedes physical actions and serves as an observable proxy for human intent. Our model is first pretrained on a large-scale egocentric human dataset to capture human intention and its synergy with action, followed by finetuning on a small set of robot and human data. During inference, the model adopts a Chain-of-Thought reasoning paradigm, sequentially predicting intention before executing the action. Extensive evaluations in simulation and real-world settings, across long-horizon and fine-grained tasks, and under few-shot and robustness benchmarks, show that our method consistently outperforms strong baselines, generalizes better, and achieves state-of-the-art performance. Project page: https://gazevla.github.io .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。