arXiv:2609.05361cs.RO2026-09

一款可表情互动的双臂人形机器人,支持手势、语音与视觉协同交互。

Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

论文配图:Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction
图 1 · 摘自论文原文
  • 双臂12自由度+可动头部,搭载实时AI处理模块
  • 手势识别准确率96%,语音识别92%,任务整体准确超90%
  • 适合人机交互、多模态智能研究的可复现原型平台

人机交互(HRI)使人类与机器人在真实环境中实现直观智能协作。本文介绍了一款为开发和集成人工智能(AI)模块于HRI任务而设计的人形机器人原型。该系统配备12个自由度(DOFs)的双臂机构和2个自由度的头部,头部带可表达情绪的LCD屏幕。所有硬件由定制控制器板控制,通过板载Jetson模块支持实时AI处理。系统集成三个AI模块:(1) 使用MediaPipe Pose和LSTM分类器进行手势识别;(2) 采用YOLO与3D定位实现物体检测;(3) 通过语音识别与大语言模型(LLM)语义解析处理语音指令。通过定位精度实验验证,平均操作误差约为1.83厘米。实验结果表明,任务整体准确率超过90%,其中手势识别达96%,语音识别达92%。结果证实该系统作为可复现、易获取的人形平台,在HRI研究与原型开发中具有有效性。

原文摘要 · Abstract (English)

Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The system features a 12 degree-of-freedom (DOFs) dual-arm mechanism and a 2 DOFs head with an expressive LCD screen to express facial emotions. All hardware components are controlled by a custom-designed controller board with real-time AI processing supported by an onboard Jetson module. The system incorporates three AI modules: (1) gesture recognition using MediaPipe Pose and an LSTM classifier, (2) object detection with YOLO and 3D localization, and (3) voice-command processing through speech recognition and large language model(LLM)-based semantic parsing. The platform is validated through experiments on positioning accuracy, with results showing average manipulation errors of approximately 1.83 cm. To demonstrate its versatility, experimental results show over 90% task accuracy, with gesture recognition reaching 96%, speech recognition reaching 92%. The results confirm the effectiveness of the proposed system as a reproducible and accessible humanoid platform for research and prototyping in HRI.

人形机器人多模态交互手势识别语音理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。