arXiv:2410.06355cs.ROcs.AI2024-10

让机器人零样本理解人话+手势,桌面上的指令解析新方法

UNCOM: Zero-shot Context-Aware Command Understanding for Tabletop Scenarios

  • 融合语音、手势和场景上下文,零样本解析自然指令
  • 在真实数据集上实现82.39%任务成功率,抗噪声与歧义能力强
  • 模块化设计提升可解释性,适合家庭服务机器人研究

本文提出UNCOM,一种新型混合框架,用于解析桌面上的人类自然语言命令。系统整合语音、手势与场景上下文信息,提取结构化可执行指令。针对家庭环境中通用人机交互的需求,UNCOM无需预定义对象模型或特定任务训练数据,支持零样本运行。基于基础与任务特定深度学习模型,实现开箱即用的语音识别、自然语言理解、手势检测与物体分割。模块化架构显式解析命令为对象-动作-目标形式,增强透明性与可解释性,便于接入符号化机器人系统。我们在TIAGo++机器人上验证系统,并在真实人机交互数据集上评估,取得82.39%的成功率,表明系统对多样性、噪声和沟通模糊性的鲁棒性。数据集、评估场景与代码均已公开,支持后续研究。

原文摘要 · Abstract (English)

This paper presents UNCOM, a novel hybrid framework for interpreting natural human commands in tabletop scenarios. The system integrates multiple sources of information -- speech, gestures, and scene context -- to extract structured, actionable instructions for robots. Addressing the need for general-purpose human-robot interaction in domestic environments, UNCOM is designed for zero-shot operation, without reliance on predefined object models or training data specific to a given task. Using foundational and task-specific deep learning models, it allows out-of-the-box speech recognition, natural language understanding, gesture detection, and object segmentation. The modular architecture enhances transparency and explainability by explicitly parsing commands into object-action-target representations, enabling integration with symbolic robotic frameworks. We demonstrate the system in a TIAGo++ robot and provide an evaluation on a real-world data set of human-robot interaction scenarios; achieving an 82.39\% success rate over our benchmark data set, highlighting the robustness of the system to diversity, noise, and communication ambiguity. The data set, evaluation scenarios, and the code are publicly available to support future research.

人机交互零样本指令理解机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。