让机器人听懂人话直接动起来,端到端实现语言控制全身动作。
LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning
- 用神经网络直接将语言指令转为全身动作,结合强化学习与策略蒸馏。
- 通过条件变分自编码器生成多样且连贯的肢体动作,支持新动作涌现。
- 在仿真和真实机器人上验证,能适应不同说法并流畅切换动作。
通用人形机器人需能与人类自然交互,融入日常生活。自然语言是实现这一目标最直观的媒介,但将语言理解转化为物理动作仍面临巨大挑战,主要源于语言与动作之间的鸿沟。本文提出一种端到端的语言驱动人形机器人全身控制策略。该方法融合强化学习与策略蒸馏,使单一神经网络可直接解析语言指令并执行相应物理动作。为提升动作多样性与组合性,引入条件变分自编码器(CVAE)结构。所提策略能根据语言输入生成敏捷、多样的全身行为,动作间过渡平滑,具备对语言变化的适应能力,并可涌现出新动作。通过大量仿真与真实实验验证了方法的有效性与泛化能力,实现了鲁棒的全身运动控制。
原文摘要 · Abstract (English)
General-purpose humanoid robots are expected to interact intuitively with humans, enabling seamless integration into daily life. Natural language provides the most accessible medium for this purpose. However, translating language into humanoid whole-body motion remains a significant challenge, primarily due to the gap between linguistic understanding and physical actions. In this work, we present an end-to-end, language-directed policy for real-world humanoid whole-body control. Our approach combines reinforcement learning with policy distillation, allowing a single neural network to interpret language commands and execute corresponding physical actions directly. To enhance motion diversity and compositionality, we incorporate a Conditional Variational Autoencoder (CVAE) structure. The resulting policy achieves agile and versatile whole-body behaviors conditioned on language inputs, with smooth transitions between various motions, enabling adaptation to linguistic variations and the emergence of novel motions. We validate the efficacy and generalizability of our method through extensive simulations and real-world experiments, demonstrating robust whole-body control. Please see our website at LangWBC.github.io for more information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。