arXiv:2503.09335cs.ROcs.AI2025-03中稿 · publication in ESW…被引 18

用语音和手势实现零样本人机交互,让机器人懂新物品

NVP-HRI: Zero Shot Natural Voice and Posture-based Human-Robot Interaction via Large Language Model

  • 结合语音与指向姿态,通过大模型理解多模态指令
  • 零样本识别新物体,实测效率比传统手势高59.2%
  • 适合老人等不擅长记命令的用户,降低使用门槛

有效的机器人交互对老龄化社会的服务机器人至关重要。现有系统仅针对预训练物体,面对新物体时表现不佳。当前基于预设手势或语言标签的交互方式对所有人(尤其是老年人)都存在记忆负担,包括难以回忆指令、掌握手部动作和学习新名称。本文提出NVP-HRI,一种融合语音与指物姿态的直观多模态交互范式。该系统利用分割一切模型(SAM)分析视觉线索与深度数据,实现精准的物体结构表征。通过预训练的SAM网络,即使无先验知识,也能实现对新物体的零样本预测。NVP-HRI还集成大语言模型(LLM),实时协调多模态指令、物体选择与场景分布,生成无碰撞轨迹。通过关键控制语法规范动作序列,降低大模型幻觉风险。在通用机器人上进行的多种真实任务评估显示,相比传统手势控制,最高提升59.2%的效率,视频演示见https://youtu.be/EbC7al2wiAc。代码与设计将公开于https://github.com/laiyuzhi/NVP-HRI.git。

原文摘要 · Abstract (English)

Effective Human-Robot Interaction (HRI) is crucial for future service robots in aging societies. Existing solutions are biased toward only well-trained objects, creating a gap when dealing with new objects. Currently, HRI systems using predefined gestures or language tokens for pretrained objects pose challenges for all individuals, especially elderly ones. These challenges include difficulties in recalling commands, memorizing hand gestures, and learning new names. This paper introduces NVP-HRI, an intuitive multi-modal HRI paradigm that combines voice commands and deictic posture. NVP-HRI utilizes the Segment Anything Model (SAM) to analyze visual cues and depth data, enabling precise structural object representation. Through a pre-trained SAM network, NVP-HRI allows interaction with new objects via zero-shot prediction, even without prior knowledge. NVP-HRI also integrates with a large language model (LLM) for multimodal commands, coordinating them with object selection and scene distribution in real time for collision-free trajectory solutions. We also regulate the action sequence with the essential control syntax to reduce LLM hallucination risks. The evaluation of diverse real-world tasks using a Universal Robot showcased up to 59.2\% efficiency improvement over traditional gesture control, as illustrated in the video https://youtu.be/EbC7al2wiAc. Our code and design will be openly available at https://github.com/laiyuzhi/NVP-HRI.git.

人机交互零样本多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。