arXiv:2505.10007cs.LGmath.OC2025-05NeurIPS被引 9

提出两种新算法,让强化学习在长期稳定表现上更可靠且样本效率更高。

Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning

  • 将鲁棒平均奖励问题转化为带锚点的分布鲁棒马尔可夫决策过程
  • 在小扰动下达到近最优样本复杂度 $\widetilde{O}(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2})$
  • 适用于机器人、医疗等需长期稳定的高可靠性场景

针对机器人、运筹和医疗等对长期性能稳定性要求高的实际应用,研究分布鲁棒(DR)平均奖励强化学习问题。提出两种算法,实现近最优样本复杂度:第一种将问题转化为分布鲁棒折扣马尔可夫决策过程;第二种引入锚定状态以稳定控制转移核在不确定性集内。假设名义MDP均匀遍历,证明两者在KL与$f_k$-散度不确定性集下,当不确定性半径足够小时,均能达到估计最优策略及鲁棒平均奖励的样本复杂度为$\widetilde{O}(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2})$。其中$\varepsilon$为目标精度,$|\mathbf{S}|$、$|\mathbf{A}|$为状态与动作空间大小,$t_{\mathrm{mix}}$为名义MDP的混合时间。这是首个针对分布鲁棒平均奖励强化学习的有限样本收敛保证。数值实验验证了算法收敛速率。

原文摘要 · Abstract (English)

Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally robust (DR) average-reward reinforcement learning. We propose two algorithms that achieve near-optimal sample complexity. The first reduces the problem to a DR discounted Markov decision process (MDP), while the second, Anchored DR Average-Reward MDP, introduces an anchoring state to stabilize the controlled transition kernels within the uncertainty set. Assuming the nominal MDP is uniformly ergodic, we prove that both algorithms attain a sample complexity of $\widetilde{O}\left(|\mathbf{S}||\mathbf{A}| t_{\mathrm{mix}}^2\varepsilon^{-2}\right)$ for estimating the optimal policy as well as the robust average reward under KL and $f_k$-divergence-based uncertainty sets, provided the uncertainty radius is sufficiently small. Here, $\varepsilon$ is the target accuracy, $|\mathbf{S}|$ and $|\mathbf{A}|$ denote the sizes of the state and action spaces, and $t_{\mathrm{mix}}$ is the mixing time of the nominal MDP. This represents the first finite-sample convergence guarantee for DR average-reward reinforcement learning. We further validate the convergence rates of our algorithms through numerical experiments.

强化学习鲁棒性样本效率马尔可夫决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。