用多源数据训练机器人通用模型,提升跨设备操作能力。
JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy

- 融合网络、人类操作视频、仿真与真实机器人数据进行预训练
- 在仿真和真实场景中均超越现有方法,尤其擅长泛化任务
- 适合需要跨机器人平台迁移技能的研究者与开发者
开放世界中的机器人自主性受限于数据多样性不足及跨机器人形态的泛化能力差。现有机器人数据集规模小、任务覆盖有限,而不同机器人形态间差异大,阻碍行为知识迁移。为此,我们提出JoyAI-RA,一种面向可泛化机器人操作的视觉-语言-动作(VLA)具身基础模型。该模型采用多源多层次预训练框架,整合网络数据、大规模第一人称人类操作视频、仿真生成轨迹及真实机器人数据。通过在异构多源数据上训练并显式统一动作空间,有效弥合人类操作与机器人控制间的形态差距,显著增强跨形态行为学习能力。在仿真与真实世界基准测试中,JoyAI-RA均优于当前最优方法,尤其在多样且需泛化的任务上表现突出。
原文摘要 · Abstract (English)
Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large differences across robot embodiments impede effective behavior knowledge transfer. To address these challenges, we propose JoyAI-RA, a vision-language-action (VLA) embodied foundation model tailored for generalizable robotic manipulation. JoyAI-RA presents a multi-source multi-level pretraining framework that integrates web data, large-scale egocentric human manipulation videos, simulation-generated trajectories, and real-robot data. Through training on heterogeneous multi-source data with explicit action-space unification, JoyAI-RA effectively bridges embodiment gaps, particularly between human manipulation and robotic control, thereby enhancing cross-embodiment behavior learning. JoyAI-RA outperforms state-of-the-art methods in both simulation and real-world benchmarks, especially on diverse tasks with generalization demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。