提出两种新算法,让强化学习在长期稳定表现上更可靠且样本效率更高。
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
- 将鲁棒平均奖励问题转化为带锚点的分布鲁棒马尔可夫决策过程
- 在小扰动下达到近最优样本复杂度 $\widetilde{O}(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2})$
- 适用于机器人、医疗等需长期稳定的高可靠性场景
针对机器人、运筹和医疗等对长期性能稳定性要求高的实际应用,研究分布鲁棒(DR)平均奖励强化学习问题。提出两种算法,实现近最优样本复杂度:第一种将问题转化为分布鲁棒折扣马尔可夫决策过程;第二种引入锚定状态以稳定控制转移核在不确定性集内。假设名义MDP均匀遍历,证明两者在KL与$f_k$-散度不确定性集下,当不确定性半径足够小时,均能达到估计最优策略及鲁棒平均奖励的样本复杂度为$\widetilde{O}(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2})$。其中$\varepsilon$为目标精度,$|\mathbf{S}|$、$|\mathbf{A}|$为状态与动作空间大小,$t_{\mathrm{mix}}$为名义MDP的混合时间。这是首个针对分布鲁棒平均奖励强化学习的有限样本收敛保证。数值实验验证了算法收敛速率。
原文摘要 · Abstract (English)
Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally robust (DR) average-reward reinforcement learning. We propose two algorithms that achieve near-optimal sample complexity. The first reduces the problem to a DR discounted Markov decision process (MDP), while the second, Anchored DR Average-Reward MDP, introduces an anchoring state to stabilize the controlled transition kernels within the uncertainty set. Assuming the nominal MDP is uniformly ergodic, we prove that both algorithms attain a sample complexity of $\widetilde{O}\left(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2}\right)$ for estimating the optimal policy as well as the robust average reward under KL and $f_k$-divergence-based uncertainty sets, provided the uncertainty radius is sufficiently small. Here, $\varepsilon$ is the target accuracy, $|\mathbf{S}|$ and $|\mathbf{A}|$ denote the sizes of the state and action spaces, and $t_{\mathrm{mix}}$ is the mixing time of the nominal MDP. This represents the first finite-sample convergence guarantee for DR average-reward reinforcement learning. We further validate the convergence rates of our algorithms through numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。