用可区分性进步引导目标选择,让强化学习更快学会多样技能。
Diversity Progress for Goal Selection in Discriminability-Motivated RL
- 根据目标可区分性的提升动态调整目标选择策略
- 比以往方法更快学会可区分的技能组合,且不出现目标分布坍缩
- 适合研究内在动机与多样化技能学习的学者
非均匀的目标选择有望提升强化学习中技能的学习效率。本文提出一种在内在动机驱动的目标条件强化学习中学习目标选择策略的方法——“多样性进展”(Diversity Progress, DP)。该方法基于观察到的目标间可区分性提升来构建学习课程。所提方法适用于以可区分性为内在奖励的智能体,其内在奖励由智能体对当前追求目标的确定性决定,从而激励智能体在无外部奖励时学习一组多样技能。实验表明,受DP驱动的智能体能比先前方法更快学习到可区分的技能集,且避免了某些已有方法中常见的目标分布坍缩问题。文章最后提出后续研究方向。
原文摘要 · Abstract (English)
Non-uniform goal selection has the potential to improve the reinforcement learning (RL) of skills over uniform-random selection. In this paper, we introduce a method for learning a goal-selection policy in intrinsically-motivated goal-conditioned RL: "Diversity Progress" (DP). The learner forms a curriculum based on observed improvement in discriminability over its set of goals. Our proposed method is applicable to the class of discriminability-motivated agents, where the intrinsic reward is computed as a function of the agent's certainty of following the true goal being pursued. This reward can motivate the agent to learn a set of diverse skills without extrinsic rewards. We demonstrate empirically that a DP-motivated agent can learn a set of distinguishable skills faster than previous approaches, and do so without suffering from a collapse of the goal distribution -- a known issue with some prior approaches. We end with plans to take this proof-of-concept forward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。