arXiv:2604.19734cs.ROcs.AI2026-04被引 10

用视觉锚定统一动作语言,让人类动作数据直接驱动人形机器人。

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

论文配图:UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
图 1 · 摘自论文原文
  • 通过视觉-动作跨模态重建,建立跨体态一致的物理意图表示。
  • 在仿真与真实场景中实现零样本任务迁移,数据效率达当前最优。
  • 适合研究人形机器人政策学习与世界建模的学者快速复现与扩展。

构建人形机器人基础模型面临机器人数据稀缺瓶颈。尽管大量第一人称人类数据可规模化获取,但因运动学差异导致跨体态迁移困难。本文提出UniT(基于视觉锚定的统一潜动作标记器),以异构运动学共享统一视觉后果为哲学基础,采用三分支交叉重建机制:动作预测视觉以锚定运动学与物理结果,视觉重建动作以过滤无关视觉干扰。同时,融合分支将净化后的模态映射至与体态无关的离散潜在空间,表征通用物理意图。在两类验证中表现优异:1)策略学习(VLA-UniT):通过预测统一标记,高效利用多样化人类数据,在人形仿真基准和真实部署中均实现最先进的数据效率与分布外泛化能力,显著实现零样本任务迁移;2)世界建模(WM-UniT):通过统一标记对齐跨体态动态,实现人类到人形机器人的直接动作迁移,确保人类数据无缝转化为增强的人形视频生成控制能力。最终,通过t-SNE可视化验证了人类与人形特征收敛于共享流形,证明了跨体态表示的高度对齐性,为从海量人类知识中蒸馏通用人形能力提供了可扩展路径。

原文摘要 · Abstract (English)

Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that establishes a unified physical language for human-to-humanoid transfer. Grounded in the philosophy that heterogeneous kinematics share universal visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch synergies these purified modalities into a shared discrete latent space of embodiment-agnostic physical intents. We validate UniT across two paradigms: 1) Policy Learning (VLA-UniT): By predicting these unified tokens, it effectively leverages diverse human data to achieve state-of-the-art data efficiency and robust out-of-distribution (OOD) generalization on both humanoid simulation benchmark and real-world deployments, notably demonstrating zero-shot task transfer. 2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it realizes direct human-to-humanoid action transfer. This alignment ensures that human data seamlessly translates into enhanced action controllability for humanoid video generation. Ultimately, by inducing a highly aligned cross-embodiment representation (empirically verified by t-SNE visualizations revealing the convergence of human and humanoid features into a shared manifold), UniT offers a scalable path to distill vast human knowledge into general-purpose humanoid capabilities.

人形机器人动作迁移视觉-动作对齐零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。