揭示强化学习中策略演化的几何规律,解释为何模型会阶段性改进。
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
- 用局部学习系数分析策略优化的几何特性
- 实验证明训练中出现高损到低损的阶段性跃迁
- 适合研究深度强化学习机制与收敛行为的人
奇异学习理论将贝叶斯学习描述为精度与复杂性之间的动态权衡,随样本量增加产生质变解。本文将该理论扩展至强化学习,证明广义后验在策略空间的集中受局部学习系数(LLC)控制,而LLC是损失函数几何的不变量。该理论预测:使用SGD的深度强化学习应从高损失的简单策略逐步演变为低损失的复杂策略。我们在一个网格世界环境中验证了这一预测,观察到训练过程中的相变表现为‘对立阶梯’——此时损失急剧下降,而LLC上升。
原文摘要 · Abstract (English)
Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as sample size increases. We extend this theory to reinforcement learning, proving that the concentration of a generalized posterior over policies is governed by the local learning coefficient (LLC), an invariant of the geometry of the regret function. This theory predicts that deep reinforcement learning with SGD should proceed from simple policies with high regret to complex policies with low regret. We verify this prediction empirically in a gridworld environment exhibiting stagewise policy development: phase transitions over training manifest as "opposing staircases" where regret decreases sharply while the LLC increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。