用双层动作对齐,让机器人从人类视频中学习抓取技能
JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

- 通过隐式和显式动作对齐,统一不同来源的动作数据
- 人类视频越多,机器人任务表现越好,未见任务也提升
- 适合想用海量人类视频训练机器人的研究者
机器人数据稀缺,通用策略需融合人类第一视角视频、仿真和真实机器人数据,但这些数据在监督信号和具身形态上差异大,动作标签缺失或不兼容。人类视频规模最大但与机器人数据差距最远,简单拼接会导致负迁移而非知识共享。我们提出 JoyAI-RA 0.5,一种视觉-语言-世界-动作(VLWA)通用框架,结合物理世界动态先验与视觉语义,通过双动作对齐实现跨数据源的操控学习扩展。隐式动作对齐从视觉变化推断潜在动作,使无动作的人类、仿真和机器人数据可指导基于潜动作条件的世界模型学习物理动态;显式对齐通过标准动作表示和相机帧相对末端执行器动作,将可靠的人类与机器人轨迹对齐至统一物理动作空间。内-外层强化学习阶段实现高效任务适配与基础策略优化。在真实 AgiBot 基准上,JoyAI-RA 在已见任务和未见变体上均表现优异。任务得分随人类第一视角预训练数据量增加持续提升,未见饱和迹象。这表明大量弱标注的人类经验可转化为可迁移的训练信号,使人类视频不仅为辅助来源,更是操控能力规模化的主要轴线。
原文摘要 · Abstract (English)
Robot data is scarce, so generalist policies need to learn from heterogeneous sources, including human egocentric video, simulation, and real robots, which differ in supervision and embodiment, with action labels missing or mutually incompatible. Human egocentric data scale best but sit farthest from robot data, and naive pooling causes negative transfer rather than knowledge sharing. We propose JoyAI-RA 0.5, a generalist Vision-Language-World-Action (VLWA) framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment. Implicit action alignment infers latent actions from visual transitions, enabling action-free human, simulation, and robot data to guide a latent-action-conditioned world model in learning physical dynamics. Explicit alignment grounds reliable human and robot trajectories in a unified physical action space through a canonical action representation and camera-frame chunk-relative end-effector actions. An inner-outer-loop reinforcement stage then pairs efficient task adaptation with foundation-policy improvement. On a real-world AgiBot benchmark, JoyAI-RA performs strongly on both seen tasks and unseen variations. The task score improves consistently as the volume of human egocentric pretraining data increases and shows no sign of plateauing at our largest scale. This suggests that abundant but weakly labeled human experience can be converted into a transferable training signal, making human video not merely a weak auxiliary source but a primary axis along which manipulation capability can be scaled. Project page can be found at https://joyai-ra-05.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。