arXiv:2511.22963cs.ROcs.AI2025-11被引 7

让人形机器人听懂自然语言指令并生成多样稳定动作

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary

  • 构建统一的动作词汇表,连接语言与物理动作
  • 两阶段训练提升新指令泛化能力,动作多样性高且稳定
  • 适合需要自然语言交互的机器人研发人员

使人形机器人能够理解自由形式的自然语言指令,是实现无缝人机交互和通用具身智能的关键一步。然而,现有方法仍受限于简单指令,或为保证物理合理性牺牲动作多样性。为此,我们提出 Humanoid-LLA——一种大型语言动作模型,可将不受限的自然语言直接转化为可用于人形机器人的全身可执行动作。该方法解决两大核心挑战:语言-动作配对数据稀缺与物理不稳定性。首先,通过学习统一的人类-人形动作词汇表,建立高层语义与物理控制之间的桥梁;其次,提出新型两阶段微调框架,先进行监督式运动思维链学习,再通过强化学习结合物理反馈优化,确保动作鲁棒性与稳定性。仿真与真实世界跨本体实验表明,Humanoid-LLA 在新语言指令上的泛化能力更强,能生成多样化且高物理保真度的动作。

原文摘要 · Abstract (English)

Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing methods remain limited, often constrained to simple instructions or forced to sacrifice motion diversity for physical plausibility. To address this gap, we present Humanoid-LLA, a Large Language Action model that translates unconstrained natural language directly into executable whole-body motions for humanoid robots. Our approach tackles two core challenges: paired language-humanoid motion data scarcity and physical instability. First, we bridge high-level language semantics with physically-grounded control by learning a unified human-humanoid motion vocabulary. Second, we introduce a novel two-stage fine-tuning framework that begins with supervised motion Chain-of-Thought learning, followed by reinforcement learning refined with physical feedback to ensure robustness and stability. Extensive evaluation in simulation and real-world cross-embodiment experiments demonstrates that Humanoid-LLA achieves superior generalization to novel language commands and diverse motion generation while maintaining high physical fidelity.

人形机器人语言动作具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。