通过组合扩散策略,用少量示范数据快速学会新动作
Composing Diffusion Policies for Few-shot Learning of Movement Trajectories
- 用扩散模型概率组合基础动作策略,拟合少样本演示数据分布
- 在多种动作上使MMD-FK误差降低超30%
- 适合需快速学习新动作的机器人场景,无需目标导向
人类能通过组合已有的身体技能完成新动作,如行走中挥棒而无需重新学习。让机器人具备这种能力对快速学习新技能至关重要。本文提出一种名为DSE(Diffusion Score Equilibrium)的新组合方法,通过概率性组合扩散策略,更好地建模少样本演示数据分布,实现非目标导向的机器人运动少样本学习。由于缺乏通用的技能与演示间误差评估指标,我们引入任务与动作空间无关的统计量——前向运动学核上的最大均值差异(MMD-FK)。实验表明,使用DSE方法后,各类技能的MMD-FK误差平均下降超过30%。进一步在真实机器人上验证,仅需5次示范即可成功教授新轨迹。
原文摘要 · Abstract (English)
Humans can perform various combinations of physical skills without having to relearn skills from scratch every single time. For example, we can swing a bat when walking without having to re-learn such a policy from scratch by composing the individual skills of walking and bat swinging. Enabling robots to combine or compose skills is essential so they can learn novel skills and tasks faster with fewer real world samples. To this end, we propose a novel compositional approach called DSE- Diffusion Score Equilibrium that enables few-shot learning for novel skills by utilizing a combination of base policy priors. Our method is based on probabilistically composing diffusion policies to better model the few-shot demonstration data-distribution than any individual policy. Our goal here is to learn robot motions few-shot and not necessarily goal oriented trajectories. Unfortunately we lack a general purpose metric to evaluate the error between a skill or motion and the provided demonstrations. Hence, we propose a probabilistic measure - Maximum Mean Discrepancy on the Forward Kinematics Kernel (MMD-FK), that is task and action space agnostic. By using our few-shot learning approach DSE, we show that we are able to achieve a reduction of over 30% in MMD-FK across skills and number of demonstrations. Moreover, we show the utility of our approach through real world experiments by teaching novel trajectories to a robot in 5 demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。