揭示神经强化学习中状态可达集的几何结构,发现其维度与动作空间相关。
Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces
- 用几何视角分析连续动作空间下策略可到达的状态集。
- 证明状态流形维度约等于动作空间维度,首次建立二者关联。
- 在多个高自由度环境验证效果,可改进复杂控制任务性能。
强化学习在连续状态与动作空间的复杂任务中取得显著进展,但多数理论研究仍局限于有限状态与动作空间。本文通过几何视角,分析基于半梯度方法训练的参数化策略所诱导的局部可达状态集。我们证明,在两层神经网络策略与演员-评论家算法下,训练动态生成的可达状态形成一个低维流形,嵌入于高维原始状态空间中。在特定条件下,该流形的维度约为动作空间维度。这是首个将状态空间几何与动作空间维度直接关联的结果。我们在四个MuJoCo环境及一个可变维度的玩具环境中验证了该上界。进一步,通过在策略与价值函数网络中引入局部流形学习层,仅替换一层网络即可学习稀疏表示,显著提升高自由度控制任务的性能。
原文摘要 · Abstract (English)
Advances in reinforcement learning (RL) have led to its successful application in complex tasks with continuous state and action spaces. Despite these advances in practice, most theoretical work pertains to finite state and action spaces. We propose building a theoretical understanding of continuous state and action spaces by employing a geometric lens to understand the locally attained set of states. The set of all parametrised policies learnt through a semi-gradient based approach induces a set of attainable states in RL. We show that the training dynamics of a two-layer neural policy induce a low dimensional manifold of attainable states embedded in the high-dimensional nominal state space trained using an actor-critic algorithm. We prove that, under certain conditions, the dimensionality of this manifold is of the order of the dimensionality of the action space. This is the first result of its kind, linking the geometry of the state space to the dimensionality of the action space. We empirically corroborate this upper bound for four MuJoCo environments and also demonstrate the results in a toy environment with varying dimensionality. We also show the applicability of this theoretical result by introducing a local manifold learning layer to the policy and value function networks to improve the performance in control environments with very high degrees of freedom by changing one layer of the neural network to learn sparse representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。