用频谱分析快速评估机器人示范数据质量,提升真实场景下模仿学习效果。
An Efficient Metric for Data Quality Measurement in Imitation Learning

- 基于轨迹功率谱密度构建无须环境交互的自动化数据评分方法。
- 低频谱值对应平滑高质量示范,实测提升任务成功率与执行流畅度。
- 适合在真实场景中对用户生成的低质量示范数据进行高效筛选。
模仿学习(IL)虽取得显著进展,但实际部署仍受分布外(OOD)场景制约。通过在部署环境中收集用户示范并微调预训练策略是应对该挑战的可行方案,然而用户示范常包含大量修正动作、振荡和突变,导致策略性能下降。现有自动数据筛选方法需依赖策略回放,计算成本高且不适用于真实部署。本文提出一种快速、高效且完全自动化的示范数据排序指标,基于示范轨迹的功率谱密度(PSD)。该方法无需策略训练、环境交互或专家标注,适用于大规模现场数据清洗。较低的PSD值表示更平滑、高质量的示范,较高的则表明轨迹存在扰动与异常。我们在两个基准模仿学习数据集上验证该方法,涵盖专家与普通用户示范,并在养老院开展用户研究,使用采集示范微调π0.5模型完成日常任务。结果表明,经PSD筛选的数据所生成的策略,在任务成功率和执行轨迹平滑性上均优于未筛选基线及两种竞争性数据排序方法。
原文摘要 · Abstract (English)
Imitation learning (IL) has seen remarkable progress, yet field deployment of IL-powered robots remains hindered by the challenge of out-of-distribution (OOD) scenarios. Fine-tuning pre-trained policies with end-user demonstrations collected in deployment environments is a promising strategy to address this challenge. However, end-user demonstrations are frequently of poor quality, characterized by excessive corrective motions, oscillations, and abrupt adjustments that degrade both learned and fine-tuned policy performance. Existing automated approaches for curating demonstration data require policy rollouts in the environment, making them computationally expensive and impractical for real-world deployment. In this paper, we propose a fast, efficient, and fully automated demonstration ranking metric based on the power spectral density (PSD) of demonstration trajectories. The PSD metric requires no policy learning, environment interaction, or expert labeling, making it well-suited for scalable, in-the-field data curation. Lower PSD values correspond to smoother, higher-quality demonstrations, while higher PSD values indicate erratic, artifact-laden trajectories. We evaluate the proposed metric on two benchmark imitation learning datasets comprising expert and lay-user demonstrations, and through a user study with older adults at a retirement facility, where collected demonstrations are used to fine-tune $\pi0.5$ \cite{intelligence2025pi_} for a daily living task. Results demonstrate that PSD-curated data yields policies with higher task success rates and smoother execution trajectories compared to uncurated baselines and two competitive data-ranking methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。