用预训练模型代替人类操作,大幅降低交互式模仿学习的负担。
Easy-IIL: Reducing Human Operational Burden in Interactive Imitation Learning via Assistant Experts
- 用现成模型作为助手专家,替代大部分人工操作。
- 仅需一次示范+关键状态干预,数据质量仍保持稳定。
- 仿真与真实实验均验证减负效果,用户体验更轻松。
交互式模仿学习(IIL)通常依赖大量人工参与离线示范和在线交互。以往研究主要关注减少被动监控的负担,而非主动操作。值得注意的是,基于模型的结构化模仿方法在低数据场景下表现接近端到端策略,且所需示范显著更少。然而,随着数据量增加,这类方法性能被端到端策略超越。基于此洞察,我们提出 Easy-IIL 框架:利用现成的模型基模仿方法作为助手专家,在多数数据收集过程中替代人工操作。人类专家仅需提供一次示范以初始化助手,并在任务临近失败的关键状态进行干预。Easy-IIL 可通过保持离线与在线数据质量,维持与主流 IIL 基线相当的性能。大量仿真与真实世界实验表明,Easy-IIL 显著降低人类操作负担,同时保持高性能。用户研究进一步证实该方法有效减轻了人类专家的主观工作负荷。
原文摘要 · Abstract (English)
Interactive Imitation Learning (IIL) typically relies on extensive human involvement for both offline demonstration and online interaction. Prior work primarily focuses on reducing human effort in passive monitoring rather than active operation. Interestingly, structured model-based imitation approaches achieve comparable performance with significantly fewer demonstrations than end-to-end imitation learning policies in the low-data regime. However, these methods are typically surpassed by end-to-end policies as the data increases. Leveraging this insight, we propose Easy-IIL, a framework that utilizes off-the-shelf model-based imitation methods as an assistant expert to replace active human operation for the majority of data collection. The human expert only provides a single demonstration to initialize the assistant expert and intervenes in critical states where the task is approaching failure. Furthermore, Easy-IIL can maintain IIL performance by preserving both offline and online data quality. Extensive simulation and real-world experiments demonstrate that Easy-IIL significantly reduces human operational burden while maintaining performance comparable to mainstream IIL baselines. User studies further confirm that Easy-IIL reduces subjective workload on the human expert. Project page: https://sites.google.com/view/easy-iil
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。