arXiv:2505.12462cs.LGstat.ML2025-05中稿 · ICML被引 2

提出无需环境模型的鲁棒强化学习算法,高效求解长期决策问题。

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

  • 基于采样器设计黑箱算法,直接估计最差情况性能。
  • 在多种不确定性模型下,样本复杂度逼近理论最优水平。
  • 适合数据稀缺场景,为实际部署提供理论保障。

在平均奖励准则下的鲁棒强化学习对长期决策至关重要,尤其当环境与训练动态不同时。然而,现有研究多集中于基于模型的方法,且仅提供渐近保证,难以在数据有限场景中应用。本文提出无模型算法 Robust Halpern Iteration (RHI),利用黑箱采样器准确估计最坏情况性能,并在生成模型设定下分析其有限样本复杂度。为实现该采样器,提出高阶多级蒙特卡洛估计器,偏差低于已有方法。进一步针对 KL 和 χ² 散度等不确定性集进行实例化,证明 RHI 可在样本复杂度为 Õ(SAℋ²/ε^(2+o(1))) 下获得 ε-最优鲁棒策略,其中 S、A 分别为状态和动作数,ℋ 为鲁棒最优跨度。该结果渐近匹配鲁棒平均奖励强化学习的最佳复杂度。

原文摘要 · Abstract (English)

Robust reinforcement learning (RL) under the average-reward criterion is essential for long-term decision-making, particularly when the environment may differ from its training dynamics. However, most existing studies focus on model-based settings and provide only asymptotic guarantees, hindering their principled understanding and practical deployment, especially in data-limited scenarios. We aim to close this gap by proposing a model-free algorithm, \textbf{Robust Halpern Iteration (RHI)}. We first design our algorithm based on a black-box sampling oracle, which can estimate the worst-case performance accurately. We then derive the finite sample complexity of RHI under the generative model setting, assuming the sampling oracle. To concretely design such an oracle, we propose a $K$-order multi-level Monte-Carlo estimator, which is shown to have a lower bias compared to prior methods. We further instantiate our design for multiple uncertainty models, including KL and $χ^2$ divergence sets, and show that our RHI algorithm achieves an $\varepsilon$-optimal robust policy with a sample complexity of $\tilde{\mathcal{O}}\left( \frac{SA\mathcal{H}^2}{\varepsilon^{(2+o(1))}}\right)$, where $S,A$ are the number of states and actions, and $\mathcal{H}$ is the robust optimal span. Our result asymptotically matches the best complexity in robust average reward RL.

强化学习鲁棒性无模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。