arXiv:2511.17351cs.LGcs.SY2025-11被引 1

提出分层强化学习的收敛性分析,证明其更新机制可稳定到博弈均衡点。

Convergence and stability of Q-learning in Hierarchical Reinforcement Learning

  • 基于随机逼近与微分方程方法,建立分层Q-learning的理论框架。
  • 证明算法在特定条件下收敛且稳定,解对应于一个可解释的博弈均衡。
  • 为分层强化学习提供理论支撑,适合研究者与系统设计者参考。

分层强化学习有望高效捕捉决策问题的时间结构并提升持续学习能力,但其理论保障落后于实践。本文提出一种封建式Q-learning方案,并研究其耦合更新的收敛与稳定性条件。通过随机逼近理论与常微分方程方法,我们给出一个定理,阐明封建式Q-learning的收敛与稳定性特性,为分层强化学习提供了原则性的分析框架。此外,我们证明更新过程收敛至一个可解释为适定博弈平衡点的解,为博弈论方法应用于分层强化学习打开新路径。最后,基于封建式Q-learning的实验验证了理论预测的结果。

原文摘要 · Abstract (English)

Hierarchical Reinforcement Learning promises, among other benefits, to efficiently capture and utilize the temporal structure of a decision-making problem and to enhance continual learning capabilities, but theoretical guarantees lag behind practice. In this paper, we propose a Feudal Q-learning scheme and investigate under which conditions its coupled updates converge and are stable. By leveraging the theory of Stochastic Approximation and the ODE method, we present a theorem stating the convergence and stability properties of Feudal Q-learning. This provides a principled convergence and stability analysis tailored to Feudal RL. Moreover, we show that the updates converge to a point that can be interpreted as an equilibrium of a suitably defined game, opening the door to game-theoretic approaches to Hierarchical RL. Lastly, experiments based on the Feudal Q-learning algorithm support the outcomes anticipated by theory.

强化学习分层智能收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。