让机器人听懂音乐和语音,实时自动选择动作。
Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control

- 音频流分音乐/语音分支,分别用指纹与语义嵌入处理
- 实现音乐段落与动作策略的动态匹配,语音直接触发技能库
- 在G1机器人上验证了真实场景下的稳定响应能力
近期人形机器人与强化学习的发展已使全身运动策略具备高度表现力。然而,多数机器人仍依赖预设序列或外部触发行为,难以适应动态环境。本文提出一种新型多模态协同框架,实现语义音频驱动的人形机器人全身控制,使机器人能实时自主选择并执行合适动作技能。系统接收连续音频流,分别处理音乐与语音输入:音乐通过音频指纹与语义嵌入识别曲目并实现时间对齐,建立音乐片段与运动策略间的动态映射;语音则映射到由模仿学习获得的离散技能库,支持直接人机交互。两种模态共享统一接口,通过强化学习控制流水线调度技能执行。我们在仿真环境及Unitree G1人形机器人上验证了方法的有效性,实现了稳健的模拟到现实迁移与一致的音频条件化策略选择。
原文摘要 · Abstract (English)
Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。