arXiv:2605.01529cs.RO2026-05

从用户演示中自动筛选高质量片段,提升机器人学习鲁棒性

Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy

论文配图:Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy
图 1 · 摘自论文原文
  • 通过自监督模型识别演示中的好坏片段,保留优质子任务
  • 用马氏距离检测低质量动作段,实测在仿真与真实场景均提升性能
  • 适合非专家用户提供杂乱演示的低数据环境机器人训练

模仿学习为机器人从人类用户获取多样化技能提供了可行框架。然而,多数算法假设可获得高质量示范,这在非专家用户的数据收集中并不现实——其示范常含无意错误。直接学习此类数据可能导致不安全策略行为,而因偶发错误丢弃整段示范则浪费宝贵数据,尤其在低数据场景下。本文提出GiB(Good-in-Bad)算法,可自动识别并剔除示范中的错误子任务,同时保留高质量部分。该方法先训练自监督模型提取潜在特征,并分配二值权重标记示范整体优劣;再建模高质量片段的潜在特征分布,利用马氏距离检测并评估低质量子任务。我们在Franka机器人上验证了该方法在模拟和真实世界多步任务中的有效性,证明使用混合质量示范时能训练出更鲁棒的策略。

原文摘要 · Abstract (English)

Imitation learning offers a promising framework for enabling robots to acquire diverse skills from human users. However, most imitation learning algorithms assume access to high-quality demonstrations an unrealistic expectation when collecting data from non-expert users, whose demonstrations often contain inadvertent errors. Naively learning from such demonstrations can result in unsafe policy behavior, while discarding entire demonstrations due to occasional mistakes wastes valuable data, especially in low-data settings. In this work, we introduce GiB (Good-in-Bad), an algorithm that automatically identifies and discards erroneous subtasks within demonstrations while preserving high-quality subtasks. The filtered data can then be used by any policy learning algorithm to train more robust policies. GiB first trains a self-supervised model to learn latent features and assigns binary weights to label each demonstration as good or bad. It then models the latent feature distribution of high-quality segments and uses the Mahalanobis distance to detect and evaluate poor-quality subtasks. We validate GiB on the Franka robot in both simulated and real-world multi-step tasks, demonstrating improved policy performance when learning from mixed-quality human demonstrations.

模仿学习机器人控制数据筛选鲁棒训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。