arXiv:2505.15418cs.LGcs.AI2025-05被引 4

用引导模型提升部分可观测环境下的强化学习效果

Guided Policy Optimization under Partial Observability

  • 双模型协同训练:引导者利用额外信息,学习者通过模仿学习
  • 理论证明性能接近直接强化学习,克服现有方法局限
  • 在带噪声和记忆任务中显著优于现有方法,适合复杂控制场景

部分可观测环境中的强化学习因不确定性而面临重大挑战。尽管模拟中提供的额外信息可提升训练效果,但如何有效利用仍是一个开放问题。为此,我们提出引导策略优化(Guided Policy Optimization, GPO)框架,该框架联合训练一个引导者与一个学习者。引导者利用特权信息进行优化,同时确保与主要通过模仿学习训练的策略保持一致。我们从理论上证明该学习机制可达到与直接强化学习相当的最优性,从而克服了现有方法的关键缺陷。实验表明,GPO在多种任务中表现优异,包括具有部分可观测性和噪声的连续控制任务,以及需要记忆能力的挑战,显著超越现有方法。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) in partially observable environments poses significant challenges due to the complexity of learning under uncertainty. While additional information, such as that available in simulations, can enhance training, effectively leveraging it remains an open problem. To address this, we introduce Guided Policy Optimization (GPO), a framework that co-trains a guider and a learner. The guider takes advantage of privileged information while ensuring alignment with the learner's policy that is primarily trained via imitation learning. We theoretically demonstrate that this learning scheme achieves optimality comparable to direct RL, thereby overcoming key limitations inherent in existing approaches. Empirical evaluations show strong performance of GPO across various tasks, including continuous control with partial observability and noise, and memory-based challenges, significantly outperforming existing methods.

强化学习部分可观测策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。