构建能直接控制人形机器人全身动作的视觉语言模型,提升协同操作能力。
OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation

- 通过全身体感操控接口收集数据,实现对机器人全部自由度的直接控制。
- 仅用不到一半演示时间,就超越两个顶尖基线模型在复杂任务中的表现。
- 适合研究人形机器人多模态控制与自主操作的开发者和研究人员。
全身体感人形机器人操作需协调其完整运动链,但多数系统将上下肢分离控制,限制了整体协作,行为类似轮式双臂平台。本文探究如何构建一个原生的视觉-语言-动作(VLA)模型,直接将语言和图像映射到人形机器人的所有自由度。我们通过三阶段的单变量实验开展系统性实证研究:全身体感操控、VLA模型设计与异构联合训练。研究发现:基于关节的全身体感接口优于仅部分暴露自由度的方案;在静态和轮式双臂平台预训练的VLA模型能意外地良好迁移到人形机器人完整动作空间;与HuMI(人形版UMI)联合训练可扩展策略至新物体和指令,无需额外全身体感演示。据此推出OpenHLM——一个开源的人形机器人全身体感操作方案。在跨越宽垂直范围的长时序挑战任务中,OpenHLM以不足一半的示范时间,超越两个先进基线(GR00T N1.6 和 $Ψ_0$)。代码、训练数据及模型权重已公开于 [https://openhlm-project.github.io/]。
原文摘要 · Abstract (English)
Whole-body humanoid loco-manipulation requires coordinating the robot's entire kinematic chain. However, most existing systems typically decouple the upper and lower bodies into separate controllers, limiting such coordination and yielding behaviors similar to those of a wheeled dual-arm platform. In this paper, we ask what it takes to build a whole-body native vision-language-action (VLA) model that maps language and pixels directly to all of the humanoid's degrees of freedom. We conduct a systematic empirical study organized as a roadmap of one-variable-at-a-time experiments across three phases: whole-body teleoperation, VLA model design, and heterogeneous co-training. Our study yields several intriguing findings: a joint-based whole-body teleoperation interface outperforms alternatives that only partially expose the humanoid's degrees of freedom; a VLA pretrained on static and wheeled dual-arm platforms transfers surprisingly well to a humanoid's full action space; and co-training with HuMI, the humanoid analog of UMI, extends the policy to new objects and instructions without additional whole-body teleoperation on those targets. Following this roadmap yields OpenHLM, an open-source recipe for whole-body humanoid loco-manipulation. In a challenging long-horizon task that spans a wide vertical range of the humanoid, OpenHLM outperforms two state-of-the-art humanoid VLA baselines (GR00T N1.6 and $Ψ_0$) using less than half the total demonstration time. Our code, training data, and model checkpoints are available at [https://openhlm-project.github.io/].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。