一个策略通用于多种机器人,大幅提升跨机械臂操作的数据效率。
One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation
- 用几何感知隐空间统一不同机器人的动作表示
- 跨机械臂联合训练使成功率提升超50%
- 仅需8次新设备演示即可达到72次训练的效果
跨机械臂操作对提升机器人操作的可扩展性、降低数据收集成本至关重要。然而,各机械臂间动作空间差异和结构不一致,给多源数据联合训练带来挑战。为此,我们提出One-Policy-Fits-All(OPFA)框架,实现多个机械臂共享单一通用策略。首先,通过3D卷积网络与Transformer学习几何感知隐动作表示(GaLR),构建跨机械臂的共享隐空间;随后设计统一的隐空间重映射解码器,从隐表示中提取特定机械臂的动作,无需针对每种机械臂进行解码器调优。该方法支持端到端的异构机械臂数据协同训练,涵盖多种夹持器与灵巧手(自由度任意)。在11种不同末端执行器上实验表明,利用异构数据联合训练显著提升策略性能:相比单源训练,成功率提升超过50%;新增仅8次演示(如某新机械臂)即可达到72次训练模型的性能水平。
原文摘要 · Abstract (English)
Cross-embodiment manipulation is crucial for enhancing the scalability of robot manipulation and reducing the high cost of data collection. However, the significant differences between embodiments, such as variations in action spaces and structural disparities, pose challenges for joint training across multiple sources of data. To address this, we propose One-Policy-Fits-All (OPFA), a framework that enables learning a single, versatile policy across multiple embodiments. We first learn a Geometry-Aware Latent Representation (GaLR), which leverages 3D convolution networks and transformers to build a shared latent action space across different embodiments. Then we design a unified latent retargeting decoder that extracts embodiment-specific actions from the latent representations, without any embodiment-specific decoder tuning. OPFA enables end-to-end co-training of data from diverse embodiments, including various grippers and dexterous hands with arbitrary degrees of freedom, significantly improving data efficiency and reducing the cost of skill transfer. We conduct extensive experiments across 11 different end-effectors. The results demonstrate that OPFA significantly improves policy performance in diverse settings by leveraging heterogeneous embodiment data. For instance, cross-embodiment co-training can improve success rates by more than 50% compared to single-source training. Moreover, by adding only a few demonstrations from a new embodiment (e.g., eight), OPFA can achieve performance comparable to that of a well-trained model with 72 demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。