用分割图融合不同机器人,让策略轻松跨硬件迁移
Shadow: Leveraging Segmentation Masks for Cross-Embodiment Policy Transfer
- 训练时用源与目标机器人的分割图合成输入,统一数据分布
- 真实机器人上成功率平均提升2倍以上,优于最强基线
- 无需多机器人共训,大幅降低数据需求,适合快速部署
机器人数据收集涉及多种硬件,未来差异将加剧。本文研究在单个机器人臂(源)上用专家轨迹训练策略,并评估其在无数据的另一机器人臂(目标)上的表现。提出名为Shadow的数据编辑方法:训练和测试时均用源与目标机器人的组合分割掩码替代实际机器人。该方式使训练与测试的数据分布高度一致,实现对新未见机器人的鲁棒策略迁移,且远比需在多种机器人上大规模共训的方法更高效。实验表明,Shadow在模拟环境中的多任务、多机器人场景以及真实机器人硬件上均有效,相比最强基线,成功率平均提升超过2倍。
原文摘要 · Abstract (English)
Data collection in robotics is spread across diverse hardware, and this variation will increase as new hardware is developed. Effective use of this growing body of data requires methods capable of learning from diverse robot embodiments. We consider the setting of training a policy using expert trajectories from a single robot arm (the source), and evaluating on a different robot arm for which no data was collected (the target). We present a data editing scheme termed Shadow, in which the robot during training and evaluation is replaced with a composite segmentation mask of the source and target robots. In this way, the input data distribution at train and test time match closely, enabling robust policy transfer to the new unseen robot while being far more data efficient than approaches that require co-training on large amounts of data from diverse embodiments. We demonstrate that an approach as simple as Shadow is effective both in simulation on varying tasks and robots, and on real robot hardware, where Shadow demonstrates an average of over 2x improvement in success rate compared to the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。