arXiv:2506.07040cs.LGcs.AI2025-06被引 4

提出高效鲁棒强化学习算法,直接从数据中学习长期最优策略。

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

  • 设计新半范数证明鲁棒贝尔曼算子严格收缩,支持快速收敛。
  • 算法达到ε精度时样本复杂度为$ ilde{ m O}(ε^{-2})$,优于传统方法。
  • 适用于模型不准确场景,适合需长期稳定决策的系统优化。

我们研究分布鲁棒无限时域平均回报马尔可夫决策过程(MDPs)的无模型方法。在污染、总变差距离和沃瑟斯坦不确定性集下,给出了Q-learning与演员-评论家算法的非渐近收敛分析。分析关键在于证明最优鲁棒贝尔曼算子关于精心设计的半范数是严格压缩的,这使得基于随机逼近的更新能以$ ilde{ m O}(ε^{-2})$依赖于目标精度学习最优鲁棒Q函数。同时建立了鲁棒TD的收敛界,常数对所有平稳策略一致,从而实现高效的数据驱动评论家估计。在此基础上,提出一种演员-评论家算法,可在$ ilde{ m O}(ε^{-2})$依赖下学习ε-最优鲁棒策略。数值模拟展示了所提算法的定性行为。结果为模型误设下的鲁棒规划提供了理论基础,并推动了直接从仿真数据构建鲁棒长期策略的无模型方法发展。

原文摘要 · Abstract (English)

We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learning and actor-critic algorithms for robust average-reward MDPs under contamination, total-variation distance, and Wasserstein uncertainty sets. A key ingredient of our analysis is showing that the optimal robust Bellman operator is a strict contraction with respect to a carefully designed semi-norm. This property enables a stochastic approximation update that learns the optimal robust $Q$-function with $\tilde{\mathcal{O}}(ε^{-2})$ dependence on the target accuracy. We also establish robust TD convergence bounds whose constants are uniform over all stationary policies, yielding an efficient data-driven routine for robust critic estimation. Building on this, we introduce an actor-critic algorithm that learns an $ε$-optimal robust policy with $\tilde{\mathcal{O}}(ε^{-2})$ dependence on the target accuracy. We provide numerical simulations to illustrate the qualitative behavior of the proposed algorithms. Our results contribute to the theoretical foundations of robust planning under model misspecification and to model-free approaches for building robust long-run policies directly from simulation data.

强化学习鲁棒控制无模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。