让机器人实时生成与说话内容匹配的手势,提升人机交互自然度。
RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

- 从原始语音中提取语义和语调特征,驱动手势生成。
- 在真实机器人上实现安全、流畅且语义对齐的实时手势响应。
- 适合研究人机交互、具身智能与机器人动画的开发者。
使类人机器人能根据人类语言实时生成同步且语义合理的手势,是实现自然人机交互的基础。然而,该任务面临三大挑战:语义丰富的数据集稀缺、模型忽略音频信号而依赖运动惯性(模态遮蔽)、以及仿真到现实的物理安全差距。我们提出RoboGesture,一种以机器人为核心的框架,协同设计数据、建模与控制,构建完整的实时交互系统。首先,建立包含300余种手势类别的RoboGesture数据集,并开发自动化流水线生成大规模无碰撞、适配机器人的音-动对。架构采用分层语义-声学对齐器,直接从原始音频标记中提取多粒度的语调与语义线索;这些线索驱动基于扩散Transformer的流匹配条件运动生成器。为保证高响应性,引入抗惯性CFG掩码,强制模型主动从音频中挖掘控制信号,避免陷入重复历史模式。最后,基于MPC的安全过滤器确保物理硬件上的实时无碰撞执行。在Unitree G1类人机器人上的实验表明,RoboGesture生成的手势更安全、更富节奏感,且语义相关性优于现有最佳方法。
原文摘要 · Abstract (English)
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the "modality eclipse" where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。