arXiv:2606.31158cs.ROcs.AI2026-06

用大模型融合语音、手势和音乐,让机器人自然执行复杂动作

LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

论文配图:LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music
图 1 · 摘自论文原文
  • 用大模型综合语音、手势和音乐节奏生成动作指令
  • 在四足机器人上实现多模态输入驱动的连贯动作执行
  • 适合研究人机交互与具身智能的开发者参考

实现自然人机交互仍是机器人领域的重大挑战。传统方法依赖预设命令,限制了机器人的表现力与适应性。本文提出一种新框架,利用大语言模型(LLM)从多模态人类输入——自然语音、手部手势和音乐/声音节拍——中合成复杂机器人动作。系统集成语音转写模型、手势识别模块及节拍检测信号处理流程,将处理后的输入通过提示模板上下文化后送入LLM。LLM基于预定义的机器人动作空间,对多源信息进行推理,生成连贯的动作序列,并通过ROS分发至四足机器人执行。该框架可解析语音语义、手势指代信息及音乐节奏线索,推动机器人向更流畅、创造性与情境感知的人机互动迈进。

原文摘要 · Abstract (English)

The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often rely on rigid, pre-programmed commands that limit the robot's expressiveness and adaptability. This paper introduces a novel framework that leverages the reasoning capabilities of Large Language Models (LLMs) to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats. Our system architecture integrates a speech transcription model, a gesture recognition module, and a signal processing pipeline for beat detection. These processed inputs are contextualized using prompt templates and fed into a LLM. The LLM, informed by a predefined robot action space, reasons over the combined inputs to generate a coherent sequence of actions. This sequence is dispatched to an action queue for execution on a quadruped robot over ROS. The framework has ability to interpret and fuse semantic commands from speech, deictic information from gestures, and rhythmic cues from music. This work represents a step towards creating robots that can interact with humans in a more fluid, creative, and context-aware manner.

人机交互多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。