arXiv:2510.00906eess.SYcs.AI2025-10

用随机可达管减少专家干预次数,提升交互式模仿学习效率

TubeDAgger: Reducing the Number of Expert Interventions with Stochastic Reach-Tubes

  • 引入随机可达管判断是否需要专家介入
  • 相比基于怀疑分类的模型,显著降低专家干预频次
  • 无需针对不同环境调参,通用性强,适合实时交互场景

交互式模仿学习在在线环境中从专家示范中训练新手策略。经典DAgger算法通过交替与环境交互和重新训练网络来构建鲁棒的新手策略。现有多种变体在判断是否允许新手执行动作或返回专家控制方面存在差异。本文提出使用随机可达管——一种常用于动态系统验证的方法——作为判断专家干预必要性的新机制。该方法无需为不同环境微调决策阈值,能有效减少专家干预次数,尤其在对比依赖怀疑分类模型的相关方法时表现更优。

原文摘要 · Abstract (English)

Interactive Imitation Learning deals with training a novice policy from expert demonstrations in an online fashion. The established DAgger algorithm trains a robust novice policy by alternating between interacting with the environment and retraining of the network. Many variants thereof exist, that differ in the method of discerning whether to allow the novice to act or return control to the expert. We propose the use of stochastic reachtubes - common in verification of dynamical systems - as a novel method for estimating the necessity of expert intervention. Our approach does not require fine-tuning of decision thresholds per environment and effectively reduces the number of expert interventions, especially when compared with related approaches that make use of a doubt classification model.

模仿学习强化学习可达性分析交互优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。