用统一的隐空间动作实现跨机械臂操作技能迁移
Latent Action Diffusion for Cross-Embodiment Manipulation
- 在隐动作空间中学习统一的动作表示,兼容不同末端执行器
- 多机器人联合训练使操作成功率提升25.3%
- 适合需要跨平台技能迁移的机器人系统开发者
端到端学习正成为机器人操作的强大范式,但受限于数据稀缺及不同机器人形态间动作空间的异质性。尤其多样化的末端执行器动作空间阻碍了跨形态学习与技能迁移。本文提出在隐动作空间中学习统一的动作表示,通过对比损失训练编码器,实现拟人手、人手与平行夹爪间的语义对齐。进一步,在不同末端执行器的操作数据上联合训练,使用该隐动作空间可使单策略控制多机器人,操作成功率最高提升25.3%,证明了在显著形态差异下仍能实现有效技能迁移。该方法为跨形态动作空间的统一提供了新路径,显著减少每种新机器人形态的数据收集需求,加速泛化,推动更高效可扩展的机器人学习。
原文摘要 · Abstract (English)
End-to-end learning is emerging as a powerful paradigm for robotic manipulation, but its effectiveness is limited by data scarcity and the heterogeneity of action spaces across robot embodiments. In particular, diverse action spaces across different end-effectors create barriers for cross-embodiment learning and skill transfer. We address this challenge through diffusion policies learned in a latent action space that unifies diverse end-effector actions. We first show that we can learn a semantically aligned latent action space for anthropomorphic robotic hands, a human hand, and a parallel jaw gripper using encoders trained with a contrastive loss. Second, we show that by using our proposed latent action space for co-training on manipulation data from different end-effectors, we can utilize a single policy for multi-robot control and obtain up to 25.3% improved manipulation success rates, indicating successful skill transfer despite a significant embodiment gap. Our approach using latent cross-embodiment policies presents a new method to unify different action spaces across embodiments, enabling efficient multi-robot control and data sharing across robot setups. This unified representation significantly reduces the need for extensive data collection for each new robot morphology, accelerates generalization across embodiments, and ultimately facilitates more scalable and efficient robotic learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。