arXiv:2410.15910cs.LGcs.AI2024-10ICLR被引 3

通过信息权重提升模仿学习中多样策略的恢复效果

Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning

  • 基于点互信息为状态动作对赋权,突出关键行为
  • 在多个数据集上实现更优的策略多样性与性能
  • 适合需要多样化决策策略的应用场景

从一组专家轨迹中恢复多样化的策略是模仿学习中的重要课题。传统方法在确定轨迹的隐式风格后,通常采用基础的行为克隆目标,对轨迹中每个状态-动作对同等对待。本文观察到,在许多场景中,行为风格仅与部分状态-动作对高度相关,因此提出一种新方法:在推断或指定轨迹隐式风格后,引入基于点互信息(PMI)的加权机制,增强标准行为克隆。该加权机制反映每个状态-动作对对风格学习的贡献度,使模型聚焦于最具代表性的行为片段。本文提供了理论支持,并通过大量实验证明该方法能有效从专家数据中恢复出更丰富多样的策略。

原文摘要 · Abstract (English)

Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse policies recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovery. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information. This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style. We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse policies from expert data.

模仿学习策略多样性点互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。