让机器人听音乐跳舞、听语音做手势,无需中间步骤直接生成动作。
Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- 用音频直接生成动作,跳过传统重建环节
- 在真实机器人上实现高保真音动同步,延迟低
- 适合需要自然互动的表演型机器人应用
人类能随音乐自然移动,但当前类人机器人缺乏表达性即兴能力,通常依赖预设动作或简单指令。传统从音频生成运动并再映射到机器人的方法需显式运动重建,导致级联误差、高延迟和音-动映射脱节。我们提出 RoboPerform,首个统一的音频到行走/舞蹈框架,可直接从音频生成音乐驱动的舞蹈与语音驱动的伴随手势。基于“运动=内容+风格”原则,将音频视为隐式风格信号,无需显式运动重建。该框架融合 ResMoE 教师策略以适应多样运动模式,以及基于扩散的学生策略实现音频风格注入。无再映射设计保证低延迟与高保真。实验验证表明,RoboPerform 在物理合理性与音动对齐方面表现优异,成功使机器人成为能响应音频的动态表演者。
原文摘要 · Abstract (English)
Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robots relies on explicit motion reconstruction, leading to cascaded errors, high latency, and disjointed acoustic-actuation mapping. We propose RoboPerform, the first unified audio-to-locomotion framework that can directly generate music-driven dance and speech-driven co-speech gestures from audio. Guided by the core principle of "motion = content + style", the framework treats audio as implicit style signals and eliminates the need for explicit motion reconstruction. RoboPerform integrates a ResMoE teacher policy for adapting to diverse motion patterns and a diffusion-based student policy for audio style injection. This retargeting-free design ensures low latency and high fidelity. Experimental validation shows that RoboPerform achieves promising results in physical plausibility and audio alignment, successfully transforming robots into responsive performers capable of reacting to audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。