让机器人同时理解语言和手势眼神,实现自然人机交互。
Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction

- 构建分层策略,融合语言与第一视角视觉、注视信息
- 仅靠短暂非语言信号即可触发机器人动作,用户负担显著降低
- 适合需要轻量交互的协作场景,如家庭服务或工业辅助
为实现自然的人机交互,机器人需理解人类不仅通过语言,还通过手势和注视等非语言信号表达意图。然而,当前机器人策略仅依赖语言指令作为唯一接口,忽视了非语言信号,使沟通负担过重。本文提出EDITH框架,通过智能眼镜实时采集人类的第一视角画面、注视轨迹和语音,并将语音转为语言指令。为处理丰富但嘈杂的多模态信号,设计分层策略:高层策略推断人类意图并生成一系列子任务,每个子任务由细粒度指令与关键帧(如指向目标物体的帧)构成,以在场景中锚定意图;低层策略执行这些子任务。在人机交互任务实验中,EDITH使机器人能响应仅持续短暂的非语言信号,且相比仅使用语言指令,显著降低了用户传达意图所需的努力。项目主页提供源代码及真实机器人演示视频。
原文摘要 · Abstract (English)
For natural human-robot interaction, a robot must understand human intent expressed not only through language but also through nonverbal signals such as gestures and gaze. However, current robot policies rely on language instructions as the sole interface for conveying intent, leaving nonverbal signals unused and placing the full burden of communication. In this work, we present EDITH, a robot framework that captures the human's nonverbal signals through continuous streams of first-person view and gaze from smart glasses, and uses them alongside language instructions as inputs to the robot policy. Our hardware system streams the human's first-person view, gaze, and speech to the robot in real time, transcribing the speech into language instructions. To handle these rich but noisy signals, we design a hierarchical policy in which a high-level policy infers the human's intent and produces a sequence of subtasks, where each subtask is represented as a fine-grained instruction paired with a keyframe that grounds the intent in the scene (e.g., the frame where the human points at the target object). A low-level policy then executes these subtasks. In our experiments on human-robot interactive tasks, EDITH enables the robot to act on the human's nonverbal signals even when intent is expressed only briefly, and significantly reduces user effort to convey intent compared to using language instructions alone. Visit our project page for source code and real-robot demo videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。