用人类操作数据训练机器人,让不同外形的机械臂都能通用。
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization
- 以人类操作为通用语言,统一不同机器人的动作空间
- 在30种机器人上用超3.5万小时数据预训练,提升跨平台泛化能力
- 适合做多形态机器人通用控制的研究者和开发者
我们提出Being-H0.5,一种面向多样化机器人平台的视觉-语言-动作(VLA)基础模型,旨在实现稳健的跨形态泛化。现有VLA模型常因形态差异大、数据少而表现不佳,我们提出以人为中心的学习范式,将人类交互轨迹视为物理交互的通用“母语”。为此,我们构建了目前最大的具身预训练方案UniHand-2.0,涵盖30种不同机器人形态的超过35,000小时多模态数据。方法上引入统一动作空间,将异构机器人控制映射至语义对齐的槽位中,使低资源机器人可从人类数据中迁移技能,高资源平台也可高效利用。基于此,设计统一序列建模与多任务预训练框架,连接人类示范与机器人执行。架构上采用混合变压器设计,结合创新的混合流(MoF)框架,分离共享运动基元与特定形态专家。为提升真实世界中的稳定性,引入流形保持门控以应对感官漂移,并提出通用异步分块机制,适配不同延迟与控制特性的平台。实验证明,Being-H0.5在模拟基准测试中表现领先,如LIBERO达到98.9%成功率,RoboCasa达53.9%,且在五种真实机器人平台上均展现出强大跨形态能力。
原文摘要 · Abstract (English)
We introduce Being-H0.5, a foundational Vision-Language-Action (VLA) model designed for robust cross-embodiment generalization across diverse robotic platforms. While existing VLAs often struggle with morphological heterogeneity and data scarcity, we propose a human-centric learning paradigm that treats human interaction traces as a universal "mother tongue" for physical interaction. To support this, we present UniHand-2.0, the largest embodied pre-training recipe to date, comprising over 35,000 hours of multimodal data across 30 distinct robotic embodiments. Our approach introduces a Unified Action Space that maps heterogeneous robot controls into semantically aligned slots, enabling low-resource robots to bootstrap skills from human data and high-resource platforms. Built upon this human-centric foundation, we design a unified sequential modeling and multi-task pre-training paradigm to bridge human demonstrations and robotic execution. Architecturally, Being-H0.5 utilizes a Mixture-of-Transformers design featuring a novel Mixture-of-Flow (MoF) framework to decouple shared motor primitives from specialized embodiment-specific experts. Finally, to make cross-embodiment policies stable in the real world, we introduce Manifold-Preserving Gating for robustness under sensory shift and Universal Async Chunking to universalize chunked control across embodiments with different latency and control profiles. We empirically demonstrate that Being-H0.5 achieves state-of-the-art results on simulated benchmarks, such as LIBERO (98.9%) and RoboCasa (53.9%), while also exhibiting strong cross-embodiment capabilities on five robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。