arXiv:2606.19935cs.AI2026-06

让机器人说话时动作更自然真实,直接生成可执行的物理动作。

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

论文配图:PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
图 1 · 摘自论文原文
  • 跳过人类动作中间表示,直接从语音生成机器人可执行的动作轨迹。
  • 在真实机器人上测试,动作与语音对齐度提升37%,运动更流畅稳定。
  • 适合需要实时交互的具身智能机器人应用,如人机协作、服务机器人。

humanoid机器人需要既富有表现力又与语音同步的肢体动作,且必须符合自身物理约束。现有方法多采用以人为中心的流程:先在人体模型(如SMPL-X)上生成动作,再映射到机器人。本工作发现该范式存在根本性的‘身体差距’:人体动作空间与机器人物理约束不匹配,导致动作迁移时出现一致性丢失和执行失败。分析表明,虽能保留粗粒度语义,但动作多样性显著压缩,语音韵律-动作同步性下降。为此,提出IK-EER框架,在映射过程中联合优化运动学可行性与语音-动作时间对齐。基于构建的机器人原生动作数据集,进一步提出PhysDrift,一种直接从语音生成可执行关节轨迹的具身感知生成框架。相比传统流程,PhysDrift在训练与推理全程保持身体一致性,并引入物理正则化以稳定运动动力学。大量实验与真实机器人部署验证,该方法显著提升语音-动作对齐度(+37%)、物理合理性、运动平滑性、推理效率及实时交互能力。

原文摘要 · Abstract (English)

Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generation pipelines are predominantly human-centric: motions are first generated in human-body representations such as SMPL-X and subsequently retargeted to humanoid robots. In this work, we identify a fundamental embodiment gap in this paradigm, where the mismatch between human motion manifolds and humanoid embodiment constraints disrupts embodiment consistency during motion transfer and physical execution. Through extensive analysis, we show that although retargeting can preserve coarse motion semantics, it significantly compresses motion diversity and weakens prosody-motion synchronization, limiting expressive humanoid behaviors. To address this problem, we first propose IK-EER, a prosody-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech-motion temporal alignment during retargeting. Building upon the curated robot-native motion dataset, we further introduce PhysDrift, an embodiment-aware co-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human-body representations. Unlike conventional human-centric pipelines, PhysDrift maintains embodiment consistency throughout both training and inference while incorporating physical regularization to stabilize robot motion dynamics. Extensive experiments and real-world humanoid deployment demonstrate that embodiment-aware robot-native generation substantially improves speech-motion alignment, physical plausibility, motion smoothness, inference efficiency, and real-time interaction capability.

人形机器人语音驱动动作生成物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。