用影响函数筛选优质示范数据,提升机器人学习效果
Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning
- 基于影响函数评估每条示范数据对验证集损失的贡献度
- 通过最大影响值和轨迹内聚合降低噪声,提升数据质量
- 适合需要高质量示范数据的机器人学习场景
从示范中学习已成为端到端机器人控制的有前景范式,尤其在大规模多样化数据集上。然而,人类遥操作收集的示范数据质量常成瓶颈:人为错误、操作限制与操作者差异引入噪声与次优行为,导致数据清洗依赖人工且主观。本文提出QoQ(Quality over Quantity),以训练样本对验证示范集损失下降的贡献定义数据质量,并利用影响函数高效估算该贡献。为适配机器人示范,提出两项关键技术:(i) 使用验证样本上的最大影响值捕捉最相关的状态-动作对;(ii) 聚合同轨迹内状态-动作对的影响得分,降低噪声并提升数据覆盖。仿真与真实世界实验表明,QoQ持续优于现有数据选择方法。
原文摘要 · Abstract (English)
Learning from demonstrations has emerged as a promising paradigm for end-to-end robot control, particularly when scaled to diverse and large datasets. However, the quality of demonstration data, often collected through human teleoperation, remains a critical bottleneck for effective data-driven robot learning. Human errors, operational constraints, and teleoperator variability introduce noise and suboptimal behaviors, making data curation essential yet largely manual and heuristic-driven. In this work, we propose Quality over Quantity (QoQ), a grounded and systematic approach to identifying high-quality data by defining data quality as the contribution of each training sample to reducing loss on validation demonstrations. To efficiently estimate this contribution, we leverage influence functions, which quantify the impact of individual training samples on model performance. We further introduce two key techniques to adapt influence functions for robot demonstrations: (i) using maximum influence across validation samples to capture the most relevant state-action pairs, and (ii) aggregating influence scores of state-action pairs within the same trajectory to reduce noise and improve data coverage. Experiments in both simulated and real-world settings show that QoQ consistently improves policy performances over prior data selection methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。