提出新理论框架,让强化学习收敛到可解释且多样化的最优策略。
Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning
- 通过渐进消减熵正则化与温度解耦,确保策略收敛
- 策略在极限下均匀采样所有最优动作,保持多样性
- 可精确估计返回分布,适合需要稳定策略的场景
为寻找最优策略,强化学习方法通常忽略所学策略的性质,导致难以预测其行为。本文提出一种理论框架,通过渐进消除熵正则化和温度解耦机制,保证收敛至特定最优策略。该策略在温度趋于零时具有可解释性并保持多样性,同时确保价值函数和回报分布的收敛。例如,所实现的策略会均匀采样所有最优动作。借助温度解耦技巧,我们设计了一种算法,可任意精度估计对应于该可解释、多样性保留最优策略的回报分布。
原文摘要 · Abstract (English)
In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even when successful, it is difficult to characterize which policies will be learned and what they will do. In this work, we present a theoretical framework for policy optimization that guarantees convergence to a particular optimal policy, via vanishing entropy regularization and a temperature decoupling gambit. Our approach realizes an interpretable, diversity-preserving optimal policy as the regularization temperature vanishes and ensures the convergence of policy derived objects--value functions and return distributions. In a particular instance of our method, for example, the realized policy samples all optimal actions uniformly. Leveraging our temperature decoupling gambit, we present an algorithm that estimates, to arbitrary accuracy, the return distribution associated to its interpretable, diversity-preserving optimal policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。