分析三种强化学习方法的收敛与稳定性,揭示其在确定性环境中的最优表现条件。
On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
- 基于马尔可夫决策过程的转移核分析,建立策略与价值函数的连续性理论
- 证明当转移核接近确定性时,能实现近最优行为,且收敛与稳定性可量化
- 适用于对算法可靠性有要求的机器人控制、游戏等序列决策场景
本文对离散时间的倒置强化学习、目标条件监督学习和在线决策变压器进行了严格的收敛与稳定性分析。这些算法在游戏与机器人任务中表现优异,但其理论基础局限于特定环境条件。本研究为通过监督学习或序列建模实现强化学习的广泛范式建立了理论框架。核心在于分析在何种环境条件下算法可识别最优解,并评估在微小噪声下的解稳定性。具体研究命令条件策略、值函数及目标达成目标在马尔可夫决策过程转移核下的连续性与渐近收敛性。结果表明,若转移核位于确定性核的足够小邻域内,则可实现近最优行为。相关量在确定性核处关于特定拓扑保持连续,无论渐近还是有限学习轮次后均成立。所提方法首次给出策略与值函数收敛与稳定性的显式估计。理论方面引入了段空间、商拓扑连续性与动力系统不动点理论等新概念。研究还包含典型环境分析与数值实验验证。
原文摘要 · Abstract (English)
This article provides a rigorous analysis of convergence and stability of Episodic Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning and Online Decision Transformers. These algorithms performed competitively across various benchmarks, from games to robotic tasks, but their theoretical understanding is limited to specific environmental conditions. This work initiates a theoretical foundation for algorithms that build on the broad paradigm of approaching reinforcement learning through supervised learning or sequence modeling. At the core of this investigation lies the analysis of conditions on the underlying environment, under which the algorithms can identify optimal solutions. We also assess whether emerging solutions remain stable in situations where the environment is subject to tiny levels of noise. Specifically, we study the continuity and asymptotic convergence of command-conditioned policies, values and the goal-reaching objective depending on the transition kernel of the underlying Markov Decision Process. We demonstrate that near-optimal behavior is achieved if the transition kernel is located in a sufficiently small neighborhood of a deterministic kernel. The mentioned quantities are continuous (with respect to a specific topology) at deterministic kernels, both asymptotically and after a finite number of learning cycles. The developed methods allow us to present the first explicit estimates on the convergence and stability of policies and values in terms of the underlying transition kernels. On the theoretical side we introduce a number of new concepts to reinforcement learning, like working in segment spaces, studying continuity in quotient topologies and the application of the fixed-point theory of dynamical systems. The theoretical study is accompanied by a detailed investigation of example environments and numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。