arXiv:2501.00785cs.RO2025-01中稿 · publication by IEE…被引 19

用语音+指向动作让机器人听懂老人指令,更自然可靠。

Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model

  • 结合语音与指向姿态,通过大模型理解指令
  • 在真实场景中准确率显著提升,抗干扰更强
  • 适合老年人或行动不便者使用,开源可复现

将人类意图转化为机器人指令对老龄化社会的服务机器人至关重要。现有基于手势或语音命令的人机交互系统因语法复杂或手语困难,对老年人不友好。为此,本文提出一种融合语音与指向姿势的多模态交互框架。视觉信息首先由目标检测模型处理,获取环境全局理解,并基于深度信息估计边界框。结合语音转文本命令与时间对齐的选定边界框,利用大语言模型生成机器人动作序列,同时施加关键控制语法约束以避免模型幻觉。系统在通用机器人UR3e上对不同复杂度的真实任务进行评估,结果表明该方法在准确性与鲁棒性方面均显著优于传统方案。为促进研究与应用,代码与设计将开源。

原文摘要 · Abstract (English)

Translating human intent into robot commands is crucial for the future of service robots in an aging society. Existing Human-Robot Interaction (HRI) systems relying on gestures or verbal commands are impractical for the elderly due to difficulties with complex syntax or sign language. To address the challenge, this paper introduces a multi-modal interaction framework that combines voice and deictic posture information to create a more natural HRI system. The visual cues are first processed by the object detection model to gain a global understanding of the environment, and then bounding boxes are estimated based on depth information. By using a large language model (LLM) with voice-to-text commands and temporally aligned selected bounding boxes, robot action sequences can be generated, while key control syntax constraints are applied to avoid potential LLM hallucination issues. The system is evaluated on real-world tasks with varying levels of complexity using a Universal Robots UR3e manipulator. Our method demonstrates significantly better performance in HRI in terms of accuracy and robustness. To benefit the research community and the general public, we will make our code and design open-source.

人机交互多模态大模型服务机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。