arXiv:2603.23755cs.LGcs.AI2026-03

提出无需复杂计算的自适应课程学习方法,提升高维环境下的训练效率。

Self Paced Gaussian Contextual Reinforcement Learning

  • 用闭式更新规则直接优化高斯上下文分布,避免昂贵的内循环计算
  • 在点质量、月球着陆等任务中表现优于或相当现有方法,尤其在隐藏上下文场景下更稳定
  • 适合高维连续、部分可观测的强化学习场景,兼具可扩展性与理论保证

课程学习通过从简单到复杂的任务序列提升强化学习效率。然而,许多自适应课程方法依赖计算成本高昂的内循环优化,限制了其在高维上下文空间中的可扩展性。本文提出自适应高斯课程学习(SPGL),通过闭式更新规则直接优化高斯上下文分布,避免了耗时的数值求解。SPGL保持了传统自适应方法的样本效率和适应能力,同时显著降低计算开销。我们提供了收敛性理论保证,并在点质量、月球着陆和接球等上下文强化学习基准上验证了该方法。实验结果表明,SPGL在隐藏上下文场景中表现更优,且上下文分布收敛更稳定,为复杂连续与部分可观测域中的课程生成提供了一种可扩展、有理论基础的替代方案。

原文摘要 · Abstract (English)

Curriculum learning improves reinforcement learning (RL) efficiency by sequencing tasks from simple to complex. However, many self-paced curriculum methods rely on computationally expensive inner-loop optimizations, limiting their scalability in high-dimensional context spaces. In this paper, we propose Self-Paced Gaussian Curriculum Learning (SPGL), a novel approach that avoids costly numerical procedures by leveraging a closed-form update rule for Gaussian context distributions. SPGL maintains the sample efficiency and adaptability of traditional self-paced methods while substantially reducing computational overhead. We provide theoretical guarantees on convergence and validate our method across several contextual RL benchmarks, including the Point Mass, Lunar Lander, and Ball Catching environments. Experimental results show that SPGL matches or outperforms existing curriculum methods, especially in hidden context scenarios, and achieves more stable context distribution convergence. Our method offers a scalable, principled alternative for curriculum generation in challenging continuous and partially observable domains.

强化学习课程学习高斯分布高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。