无需监督,通过状态转移建模和在线行为对齐实现更强泛化能力的模仿学习
Towards Generalisable Imitation Learning Through Conditioned Transition Estimation and Online Behaviour Alignment
- 先估计教师真实动作,再通过在线对齐优化策略
- 在5个环境中超越教师和其他方法,标准差最小
- 适合需要高鲁棒性的实际场景部署
当前基于观察的模仿学习(ILfO)方法虽有进展,但仍存在依赖动作监督、假设状态仅有单一最优动作、机械复制教师动作等问题。本文提出无监督基于观察的模仿学习(UfO),通过两阶段训练:首先估计教师在观测状态转移中的真实动作,然后通过调整智能体轨迹与教师轨迹对齐来进一步优化策略。在五个常用环境中的实验表明,UfO不仅优于教师及所有其他ILfO方法,且标准差最小,表明其在未见场景中具有更强的泛化能力。
原文摘要 · Abstract (English)
State-of-the-art imitation learning from observation methods (ILfO) have recently made significant progress, but they still have some limitations: they need action-based supervised optimisation, assume that states have a single optimal action, and tend to apply teacher actions without full consideration of the actual environment state. While the truth may be out there in observed trajectories, existing methods struggle to extract it without supervision. In this work, we propose Unsupervised Imitation Learning from Observation (UfO) that addresses all of these limitations. UfO learns a policy through a two-stage process, in which the agent first obtains an approximation of the teacher's true actions in the observed state transitions, and then refines the learned policy further by adjusting agent trajectories to closely align them with the teacher's. Experiments we conducted in five widely used environments show that UfO not only outperforms the teacher and all other ILfO methods but also displays the smallest standard deviation. This reduction in standard deviation indicates better generalisation in unseen scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。