arXiv:2410.12773cs.ROcs.AI2024-10CoRL被引 38

用语言描述生成拟人化机器人全身动作,让机器人听懂指令自然行动

Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

  • 基于人类动作数据初始化,用视觉语言模型理解语义并优化动作
  • 在仿真和真实机器人上验证,动作自然且与语言描述高度对齐
  • 适合需要人机自然交互的智能服务机器人场景

类人机器人因其类人形态,有望无缝融入人类环境。实现与人类共存与协作的关键在于理解自然语言并表现出类人行为。本文聚焦于从语言描述生成多样化的类人机器人全身运动。我们利用大规模人类动作数据集中的动作先验初始化机器人动作,并借助视觉语言模型(VLM)的常识推理能力对动作进行编辑与优化。所提方法可生成自然、富有表现力且与文本对齐的类人机器人动作,已在仿真与真实机器人实验中得到验证。更多视频展示见 https://ut-austin-rpl.github.io/Harmon/。

原文摘要 · Abstract (English)

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generating diverse whole-body motions for humanoid robots from language descriptions. We leverage human motion priors from extensive human motion datasets to initialize humanoid motions and employ the commonsense reasoning capabilities of Vision Language Models (VLMs) to edit and refine these motions. Our approach demonstrates the capability to produce natural, expressive, and text-aligned humanoid motions, validated through both simulated and real-world experiments. More videos can be found at https://ut-austin-rpl.github.io/Harmon/.

人形机器人语言驱动动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。