用眼神和说话结合让机器人更懂人话
SemanticScanpath: Combining Gaze and Speech for Situated Human-Robot Interaction Using LLMs
- 把用户眼神轨迹转成文本,与语音一起输入大模型
- 在多个任务中准确理解模糊指令,正确率显著提升
- 适合做智能交互机器人,尤其需要理解非语言信号的场景
大型语言模型(LLMs)显著提升了社交机器人的对话能力。然而,要实现自然流畅的人机交互,机器人需能将模糊或不完整的口头表达与当前物理情境及用户的非语言意图(如指向性凝视)关联起来。本文提出一种融合语音与凝视的表征方法,使LLMs具备更强的情境感知能力,准确解析模糊请求。该方法基于用户凝视轨迹(scanpath)的文本语义转换,并结合口语请求,展现LLMs对凝视行为的推理能力,可稳健忽略无关注视或干扰性视线。我们在多个任务和两种场景下验证了系统性能,结果表明其通用性和准确性均优于对照组。最后,我们实现了机器人平台上的闭环应用,完成从请求理解到执行的完整流程。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have substantially improved the conversational capabilities of social robots. Nevertheless, for an intuitive and fluent human-robot interaction, robots should be able to ground the conversation by relating ambiguous or underspecified spoken utterances to the current physical situation and to the intents expressed nonverbally by the user, such as through referential gaze. Here, we propose a representation that integrates speech and gaze to enable LLMs to achieve higher situated awareness and correctly resolve ambiguous requests. Our approach relies on a text-based semantic translation of the scanpath produced by the user, along with the verbal requests. It demonstrates LLMs' capabilities to reason about gaze behavior, robustly ignoring spurious glances or irrelevant objects. We validate the system across multiple tasks and two scenarios, showing its superior generality and accuracy compared to control conditions. We demonstrate an implementation on a robotic platform, closing the loop from request interpretation to execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。