arXiv:2603.12263cs.RO2026-03被引 26

用人类视频+机器人数据训练通用人形机器人,效果远超传统方法。

$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

  • 分阶段训练:先用人类视角视频学通用视觉动作,再用机器人数据学精确控制。
  • 仅用800小时人类视频和30小时机器人数据,成功率比基线高40%以上。
  • 开源完整训练框架与推理引擎,适合机器人研发者和研究者使用。

我们提出 $Ψ_0$(Psi-Zero),一个开放的基础模型,用于解决复杂的人形机器人运动与操作任务。现有方法常联合训练大量人类与人形机器人数据,但因两者在运动结构和动作模式上存在根本差异,导致数据效率与模型性能均不理想。为此,我们提出解耦学习策略:首先在大规模人类第一视角视频上自回归预训练视觉-动作模型(VLM)以获得通用表示;随后在高质量人形机器人数据上微调基于流的行动专家模型,实现精准关节控制。研究发现,相较于依赖嘈杂网络视频或跨体态机器人数据集的方法,使用高质量人类操作视频预训练并结合真实世界人形轨迹微调,可显著提升性能。实验证明,仅需约800小时人类视频与30小时真实机器人数据,$Ψ_0$ 在多任务上整体成功率超越基线40%以上,而后者使用超过10倍的数据量。我们将开源整个生态系统,包括数据处理流程、基础模型及实时动作推理引擎。

原文摘要 · Abstract (English)

We introduce $Ψ_0$ (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamental problem by co-training on large and diverse human and humanoid data, we argue that this strategy is suboptimal due to the fundamental kinematic and motion disparities between humans and humanoid robots. Therefore, data efficiency and model performance remain unsatisfactory despite the considerable data volume. To address this challenge, \ours\;decouples the learning process to maximize the utility of heterogeneous data sources. Specifically, we propose a staged training paradigm with different learning objectives: First, we autoregressively pre-train a VLM backbone on large-scale egocentric human videos to acquire generalizable visual-action representations. Then, we post-train a flow-based action expert on high-quality humanoid robot data to learn precise robot joint control. Our research further identifies a critical yet often overlooked data recipe: in contrast to approaches that scale with noisy Internet clips or heterogeneous cross-embodiment robot datasets, we demonstrate that pre-training on high-quality egocentric human manipulation data followed by post-training on domain-specific real-world humanoid trajectories yields superior performance. Extensive real-world experiments demonstrate that \ours\ achieves the best performance using only about 800 hours of human video data and 30 hours of real-world robot data, outperforming baselines pre-trained on more than 10$\times$ as much data by over 40\% in overall success rate across multiple tasks. We will open-source the entire ecosystem to the community, including a data processing and training pipeline, a humanoid foundation model, and a real-time action inference engine.

人形机器人基础模型动作生成数据解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。