提出概率化课程学习方法,自动为强化学习设计目标
Probabilistic Curriculum Learning for Goal-Based Reinforcement Learning
- 基于概率模型动态生成适合的训练目标
- 在连续控制与导航任务中提升学习效率与成功率
- 适合需要自适应训练路径的智能体系统
强化学习通过最大化奖励信号指导智能体与环境交互,近年来取得显著进展,得益于深度Q学习、确定性策略梯度、近端策略优化、信任域策略优化及软动作-评论家等算法以及GPU/TPU等专用计算资源的发展。引入目标以支持多模态策略成为重要研究方向,常通过分层或课程强化学习实现。此类方法将复杂行为分解为更简单的子任务,类比人类逐步学习技能(如先学会跑再学走,先学算术再学微积分)。然而,目标的完全自动化生成仍是一个开放挑战。本文提出一种新的概率化课程学习算法,用于在连续控制与导航任务中为强化学习智能体自动建议目标。
原文摘要 · Abstract (English)
Reinforcement learning (RL) -- algorithms that teach artificial agents to interact with environments by maximising reward signals -- has achieved significant success in recent years. These successes have been facilitated by advances in algorithms (e.g., deep Q-learning, deep deterministic policy gradients, proximal policy optimisation, trust region policy optimisation, and soft actor-critic) and specialised computational resources such as GPUs and TPUs. One promising research direction involves introducing goals to allow multimodal policies, commonly through hierarchical or curriculum reinforcement learning. These methods systematically decompose complex behaviours into simpler sub-tasks, analogous to how humans progressively learn skills (e.g. we learn to run before we walk, or we learn arithmetic before calculus). However, fully automating goal creation remains an open challenge. We present a novel probabilistic curriculum learning algorithm to suggest goals for reinforcement learning agents in continuous control and navigation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。