arXiv:2603.12324cs.LGcs.AI2026-03中稿 · ICLR

用热力学原理设计强化学习课程,让训练更高效。

Thermodynamics of Reinforcement Learning Curricula

  • 将任务视为流形上的坐标,用几何路径优化学习顺序。
  • 最小化额外功可得最优课程,对应任务空间测地线。
  • 提出MEW算法,自动规划最大熵RL的温度退火策略。

统计力学与机器学习的联系多次带来突破,启发了优化、泛化和表征学习的理解。本文延续这一传统,利用非平衡热力学成果形式化强化学习中的课程学习。我们提出一个几何框架,将奖励参数视为任务流形上的坐标。通过最小化超额热力学功,发现最优课程对应该任务空间中的测地线。作为应用,我们提出“MEW”(最小超额功)算法,用于生成最大熵强化学习中温度退火的合理调度方案。

原文摘要 · Abstract (English)

Connections between statistical mechanics and machine learning have repeatedly proven fruitful, providing insight into optimization, generalization, and representation learning. In this work, we follow this tradition by leveraging results from non-equilibrium thermodynamics to formalize curriculum learning in reinforcement learning (RL). In particular, we propose a geometric framework for RL by interpreting reward parameters as coordinates on a task manifold. We show that, by minimizing the excess thermodynamic work, optimal curricula correspond to geodesics in this task space. As an application of this framework, we provide an algorithm, "MEW" (Minimum Excess Work), to derive a principled schedule for temperature annealing in maximum-entropy RL.

强化学习课程学习热力学最大熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。