提出S-矩形分布鲁棒强化学习的近优样本复杂度,提升实际应用鲁棒性。
Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning
- 基于分歧度设计S-矩形对抗模型,更贴合真实场景分布差异。
- 理论证明样本复杂度达近最优,对状态、动作空间和精度依赖最优。
- 适用于需高鲁棒性的工业控制等现实强化学习任务。
分布鲁棒强化学习(DR-RL)作为一种应对训练与测试环境差异的系统方法,近年来受到广泛关注。为平衡鲁棒性、保守性和计算可追踪性,文献引入了SA-矩形和S-矩形对抗模型。尽管现有统计分析多集中于算法更简单的SA-矩形模型,但S-矩形模型在许多真实应用场景中更准确地刻画了分布差异,常能生成更有效的随机策略。本文研究基于分歧度的S-矩形DR-RL的经验值迭代算法,建立了近最优的样本复杂度界:$ ilde{O}(| ext{S}|| ext{A}|(1-γ)^{-4}^{-2})$,其中$$为目标精度,$| ext{S}|$、$| ext{A}|$为状态与动作空间大小,$γ$为折扣因子。据我们所知,这是首个在$| ext{S}|$、$| ext{A}|$和$$上同时达到最优依赖关系的分歧度S-矩形模型分析结果。通过在鲁棒库存控制问题和理论最坏情况下的数值实验,验证了该理论依赖关系,展示了所提算法的快速学习性能。
原文摘要 · Abstract (English)
Distributionally robust reinforcement learning (DR-RL) has recently gained significant attention as a principled approach that addresses discrepancies between training and testing environments. To balance robustness, conservatism, and computational traceability, the literature has introduced DR-RL models with SA-rectangular and S-rectangular adversaries. While most existing statistical analyses focus on SA-rectangular models, owing to their algorithmic simplicity and the optimality of deterministic policies, S-rectangular models more accurately capture distributional discrepancies in many real-world applications and often yield more effective robust randomized policies. In this paper, we study the empirical value iteration algorithm for divergence-based S-rectangular DR-RL and establish near-optimal sample complexity bounds of $\widetilde{O}(|\mathcal{S}||\mathcal{A}|(1-γ)^{-4}\varepsilon^{-2})$, where $\varepsilon$ is the target accuracy, $|\mathcal{S}|$ and $|\mathcal{A}|$ denote the cardinalities of the state and action spaces, and $γ$ is the discount factor. To the best of our knowledge, these are the first sample complexity results for divergence-based S-rectangular models that achieve optimal dependence on $|\mathcal{S}|$, $|\mathcal{A}|$, and $\varepsilon$ simultaneously. We further validate this theoretical dependence through numerical experiments on a robust inventory control problem and a theoretical worst-case example, demonstrating the fast learning performance of our proposed algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。