同时学习专家行为与错误示范,提升离线模仿学习效果
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
- 通过对比专家与错误行为的数据分布优化策略
- 在专家数据占优时目标函数为凸,训练更稳定
- 无需对抗训练,适合有正负示范的场景
离线模仿学习通常仅利用专家示范和未标注示范,却忽略了显式错误行为中的有用信号。本文研究基于对比行为的离线模仿学习,数据集包含专家与错误示范。我们提出一种新公式,优化专家与错误数据在状态-动作访问分布上的KL散度差值。尽管该目标是DC(凸差)规划,但当专家示范占比超过错误示范时,其变为凸问题,从而获得一个实用且稳定的非对抗性训练目标。本方法避免了对抗训练,统一处理正负示范。在标准离线模仿学习基准上的大量实验表明,该方法持续优于现有最先进基线。
原文摘要 · Abstract (English)
Offline imitation learning typically learns from expert and unlabeled demonstrations, yet often overlooks the valuable signal in explicitly undesirable behaviors. In this work, we study offline imitation learning from contrasting behaviors, where the dataset contains both expert and undesirable demonstrations. We propose a novel formulation that optimizes a difference of KL divergences over the state-action visitation distributions of expert and undesirable (or bad) data. Although the resulting objective is a DC (Difference-of-Convex) program, we prove that it becomes convex when expert demonstrations outweigh undesirable demonstrations, enabling a practical and stable non-adversarial training objective. Our method avoids adversarial training and handles both positive and negative demonstrations in a unified framework. Extensive experiments on standard offline imitation learning benchmarks demonstrate that our approach consistently outperforms state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。