arXiv:2504.09532cs.ROcs.AI2025-04被引 3

用多模态大模型实现人形机器人零样本协同运动与操作

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

  • 通过动作链推理分解指令为运动与操作步骤
  • 在两个真实人形机器人上实现零样本泛化,性能显著优于基线
  • 适合研究具身智能、机器人协作与自然语言控制的学者

人形机器人协同运动与操作(loco-manipulation)融合全身运动与灵巧操作,仍是机器人领域的核心挑战。除全身协调与平衡外,理解人类指令并转化为连贯的具身动作序列尤为困难。近期基础模型虽提供可迁移的多模态表征与推理能力,但现有工作多局限于运动或操作单独处理,难以应用于人形机器人。本文提出Humanoid-COA,首个将基础模型推理与具身动作链(CoA)机制结合的人形机器人零样本协同运动与操作框架。在感知-推理-动作范式中,核心贡献在于推理阶段:通过可操作性分析、空间推理与全身动作推理,将高层人类指令分解为结构化的运动与操作基元序列。在Unitree H1-2与G1两款人形机器人上,在开放测试区与公寓环境中的大量实验表明,该框架在操作、运动及协同任务中均显著超越先前基线,具备对长时程与非结构化场景的鲁棒泛化能力。

原文摘要 · Abstract (English)

Humanoid loco-manipulation, which integrates whole-body locomotion with dexterous manipulation, remains a fundamental challenge in robotics. Beyond whole-body coordination and balance, a central difficulty lies in understanding human instructions and translating them into coherent sequences of embodied actions. Recent advances in foundation models provide transferable multimodal representations and reasoning capabilities, yet existing efforts remain largely restricted to either locomotion or manipulation in isolation, with limited applicability to humanoid settings. In this paper, we propose Humanoid-COA, the first humanoid agent framework that integrates foundation model reasoning with an Embodied Chain-of-Action (CoA) mechanism for zero-shot loco-manipulation. Within the perception--reasoning--action paradigm, our key contribution lies in the reasoning stage, where the proposed CoA mechanism decomposes high-level human instructions into structured sequences of locomotion and manipulation primitives through affordance analysis, spatial inference, and whole-body action reasoning. Extensive experiments on two humanoid robots, Unitree H1-2 and G1, in both an open test area and an apartment environment, demonstrate that our framework substantially outperforms prior baselines across manipulation, locomotion, and loco-manipulation tasks, achieving robust generalization to long-horizon and unstructured scenarios. Project page: https://humanoid-coa.github.io/

人形机器人具身智能零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。