用互信息评估机器人示范数据质量,提升模仿学习效果
Robot Data Curation with Mutual Information Estimators
- 基于状态与动作的互信息,量化单条轨迹贡献度
- 在仿真与真实场景中,能按人工评分区分数据优劣
- 过滤低质数据后,策略性能提升5-10%,适用于实机部署
模仿学习策略的性能高度依赖训练数据集的质量。尽管机器人领域示范数据量显著增加,但对数据质量的评估研究仍不足。本文提出一种方法,通过估计轨迹对整体数据集中状态与动作间互信息的平均贡献,衡量单条示范在动作多样性与可预测性方面的相对质量。该方法基于状态和动作的简单VAE嵌入,结合k近邻互信息估计,在数据量有限的机器人场景中表现良好。实验表明,该方法能在多种仿真与真实环境基准上根据人类专家评分有效划分数据质量。使用该方法筛选的数据训练策略后,RoboMimic性能提升5-10%,在真实ALOHA和Franka平台上也取得更好表现。
原文摘要 · Abstract (English)
The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, little work has sought to assess the quality of said data despite mounting evidence of its importance in other areas such as vision and language. In this work, we take a critical step towards addressing the data quality in robotics. Given a dataset of demonstrations, we aim to estimate the relative quality of individual demonstrations in terms of both action diversity and predictability. To do so, we estimate the average contribution of a trajectory towards the mutual information between states and actions in the entire dataset, which captures both the entropy of the marginal action distribution and the state-conditioned action entropy. Though commonly used mutual information estimators require vast amounts of data often beyond the scale available in robotics, we introduce a novel technique based on k-nearest neighbor estimates of mutual information on top of simple VAE embeddings of states and actions. Empirically, we demonstrate that our approach is able to partition demonstration datasets by quality according to human expert scores across a diverse set of benchmarks spanning simulation and real world environments. Moreover, training policies based on data filtered by our method leads to a 5-10% improvement in RoboMimic and better performance on real ALOHA and Franka setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。