arXiv:2602.20219cs.ROcs.AI2026-02被引 2

用语音+视觉让机器人更懂人话,精准操控机械臂

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

  • 融合语音识别、视觉理解与模糊逻辑,实现多模态交互
  • 在普通电脑上达成75%指令执行准确率
  • 适合想做自然人机协作的开发者和研究者

准确理解人类意图是人机交互中的核心挑战,也是实现自然协同的关键。本文提出一种新型多模态人机交互框架,结合先进视觉-语言模型、语音处理与模糊逻辑,实现对 Dobot Magician 机械臂的精准自适应控制。系统集成 Florence-2 用于目标检测,Llama 3.1 实现自然语言理解,Whisper 完成语音识别,使用户可通过语音命令无缝操控物体。通过联合处理场景感知与动作规划,提升指令解析与执行的可靠性。在消费级硬件上的实验表明,指令执行准确率达75%,验证了系统的鲁棒性与可扩展性。该架构为未来人机交互研究提供了灵活、可扩展的基础,推动通过紧密耦合的语音与视觉-语言处理实现更自然的协作。

原文摘要 · Abstract (English)

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI framework that combines advanced vision-language models, speech processing, and fuzzy logic to enable precise and adaptive control of a Dobot Magician robotic arm. The proposed system integrates Florence-2 for object detection, Llama 3.1 for natural language understanding, and Whisper for speech recognition, providing users with a seamless and intuitive interface for object manipulation through spoken commands. By jointly addressing scene perception and action planning, the approach enhances the reliability of command interpretation and execution. Experimental evaluations conducted on consumer-grade hardware demonstrate a command execution accuracy of 75\%, highlighting both the robustness and adaptability of the system. Beyond its current performance, the proposed architecture serves as a flexible and extensible foundation for future HRI research, offering a practical pathway toward more sophisticated and natural human-robot collaboration through tightly coupled speech and vision-language processing.

人机交互多模态语音理解机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。