机器人通过在线经验自动筛选低质量演示,提升模仿学习效果。
Curating Demonstrations using Online Experience
- 用策略表现区分成功与失败的轨迹,构建分类器筛选演示数据。
- 真实世界与仿真中性能提升15%-35%,显著优于全量数据训练。
- 无需人工标注,适合大规模演示数据清洗与自主机器人系统。
许多机器人演示数据集包含质量参差不齐的多样化演示,这种异质性虽有益于策略预训练,但在最终模仿学习中可能降低机器人性能。部分策略可靠性差或数据中代表性不足,导致测试时采样后表现不佳。且这些低质策略常难以人工识别,数据筛选耗时费力。我们提出基于在线机器人经验的自筛选方法(Demo-SCORE):训练并交叉验证分类器以区分成功与失败的策略轨迹,并用其过滤异质演示数据集。仿真与真实世界实验表明,Demo-SCORE可有效识别低效演示,无需人工标注。相比使用全部原始演示训练的基础策略,其最终策略成功率提升超过15%-35%。
原文摘要 · Abstract (English)
Many robot demonstration datasets contain heterogeneous demonstrations of varying quality. This heterogeneity may benefit policy pre-training, but can hinder robot performance when used with a final imitation learning objective. In particular, some strategies in the data may be less reliable than others or may be underrepresented in the data, leading to poor performance when such strategies are sampled at test time. Moreover, such unreliable or underrepresented strategies can be difficult even for people to discern, and sifting through demonstration datasets is time-consuming and costly. On the other hand, policy performance when trained on such demonstrations can reflect the reliability of different strategies. We thus propose for robots to self-curate based on online robot experience (Demo-SCORE). More specifically, we train and cross-validate a classifier to discern successful policy roll-outs from unsuccessful ones and use the classifier to filter heterogeneous demonstration datasets. Our experiments in simulation and the real world show that Demo-SCORE can effectively identify suboptimal demonstrations without manual curation. Notably, Demo-SCORE achieves over 15-35% higher absolute success rate in the resulting policy compared to the base policy trained with all original demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。