arXiv:2412.07057stat.MLcs.LG2024-12NeurIPS

状态级交互标注可超越行为克隆,混合方法更高效。

Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning

  • 以状态为单位计算成本,交互式算法优于仅用离线数据的模仿学习。
  • 新算法Stagger在低标注成本下理论性能超越行为克隆,混合方法更稳健。
  • 实验验证:少量交互成本即可显著提升效果,适合实际应用。

模仿学习通过专家示范学习序列决策策略,依赖离线演示、交互标注或两者结合。以往研究认为,当按轨迹计费时,仅使用离线数据的行为克隆(BC)无法被一般性改进,限制了交互方法如DAgger的应用。本文重新审视该结论,证明当按状态计费时,使用交互标注的算法可严格优于BC。具体而言:(1) 提出单样本每轮的Stagger算法,其在低恢复成本条件下理论性能超越BC;(2) 首次研究混合模仿学习,提出Warm Stagger,其性能接近仅用任一数据源的效果;并给出一个马尔可夫决策过程实例,显示Warm Stagger在缓解累积误差与冷启动问题上具有显著优势;(3) 在MuJoCo连续控制任务上的实验表明,即使交互标注成本略高,交互与混合方法也持续优于BC。本工作首次揭示状态级交互标注与混合反馈在模仿学习中的价值。

原文摘要 · Abstract (English)

Imitation learning (IL) is a paradigm for learning sequential decision making policies from experts, leveraging offline demonstrations, interactive annotations, or both. Recent advances show that when annotation cost is tallied per trajectory, Behavior Cloning (BC) which relies solely on offline demonstrations cannot be improved in general, leaving limited conditions for interactive methods such as DAgger to help. We revisit this conclusion and prove that when the annotation cost is measured per state, algorithms using interactive annotations can provably outperform BC. Specifically: (1) we show that Stagger, a one sample per round variant of DAgger, provably beats BC under low recovery cost settings; (2) we initiate the study of hybrid IL where the agent learns from offline demonstrations and interactive annotations. We propose Warm Stagger whose learning guarantee is not much worse than using either data source alone. Furthermore, motivated by compounding error and cold start problem in imitation learning practice, we give an MDP example in which Warm Stagger has significant better annotation cost; (3) experiments on MuJoCo continuous control tasks confirm that, with modest cost ratio between interactive and offline annotations, interactive and hybrid approaches consistently outperform BC. To the best of our knowledge, our work is the first to highlight the benefit of state wise interactive annotation and hybrid feedback in imitation learning.

模仿学习交互标注混合学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。