用影响函数筛选机器人示范数据,提升策略性能
CUPID: Curating Data your Robot Loves with Influence Functions
- 基于影响函数理论评估每条示范数据对策略表现的影响
- 仅用33%精选数据即可达到顶尖扩散模型效果
- 适合需要高效数据筛选的机器人学习研究者
在机器人模仿学习中,策略性能与示范数据的质量和构成密切相关。然而,准确理解单条示范如何影响下游结果(如闭环任务成功或失败)仍是一大挑战。我们提出CUPID,一种基于新型影响函数理论的机器人数据筛选方法。给定一组评估轨迹,CUPID估计每条训练示范对策略期望回报的影响,从而根据其对闭环性能的作用进行排序与选择。利用CUPID,可实现:1)剔除损害策略性能的示范数据;2)挑选最能提升策略的新采集轨迹。大量仿真与硬件实验表明,该方法能持续识别驱动测试时表现的关键数据。例如,在模拟的RoboMimic基准上,使用不足33%的精选数据即可训练出当前最优的扩散策略,硬件实验也观察到类似增益。此外,硬件实验显示该方法能识别分布外下的稳健策略、分离虚假相关性,并增强通用机器人策略的后期训练。视频与代码已公开:https://cupid-curation.github.io。
原文摘要 · Abstract (English)
In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data. Yet, developing a precise understanding of how individual demonstrations contribute to downstream outcomes - such as closed-loop task success or failure - remains a persistent challenge. We propose CUPID, a robot data curation method based on a novel influence function-theoretic formulation for imitation learning policies. Given a set of evaluation rollouts, CUPID estimates the influence of each training demonstration on the policy's expected return. This enables ranking and selection of demonstrations according to their impact on the policy's closed-loop performance. We use CUPID to curate data by 1) filtering out training demonstrations that harm policy performance and 2) subselecting newly collected trajectories that will most improve the policy. Extensive simulated and hardware experiments show that our approach consistently identifies which data drives test-time performance. For example, training with less than 33% of curated data can yield state-of-the-art diffusion policies on the simulated RoboMimic benchmark, with similar gains observed in hardware. Furthermore, hardware experiments show that our method can identify robust strategies under distribution shift, isolate spurious correlations, and even enhance the post-training of generalist robot policies. Videos and code are made available at: https://cupid-curation.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。