arXiv:2411.07514cs.LGstat.ML2024-11被引 1

提出新算法解决非马尔可夫强化学习的鲁棒离线学习问题。

Robust Offline Reinforcement Learning for Non-Markovian Decision Processes

  • 基于低秩结构设计数据提炼与下置信界机制。
  • 仅需 O(1/ε²) 样本即可找到 ε-最优鲁棒策略。
  • 适用于无特定结构的模型,适合实际部署场景。

分布鲁棒离线强化学习旨在利用从标准模型收集的离线数据,在不确定性集所定义的最差环境中找到最优策略。尽管现有研究聚焦于马尔可夫决策过程(MDPs),但针对非马尔可夫决策过程的鲁棒离线强化学习仍局限于已知转移结构的规划问题。本文研究了非马尔可夫决策过程的鲁棒离线学习问题。当标准模型具有低秩结构时,提出一种新算法,包含新颖的数据集提炼方法和针对不同不确定性集的鲁棒值下置信界设计,并推导出非马尔可夫强化学习中鲁棒值的新对偶形式,提升算法实用性。通过引入一种针对离线低秩非马尔可夫决策过程的新类型-I集中系数,证明该算法在使用 O(1/ε²) 离线样本下可找到 ε-最优鲁棒策略。此外,还将算法扩展至标准模型无特定结构的情形,结合新的类型-II集中系数,在所有类型的不确定性集下仍保持多项式样本效率。

原文摘要 · Abstract (English)

Distributionally robust offline reinforcement learning (RL) aims to find a policy that performs the best under the worst environment within an uncertainty set using an offline dataset collected from a nominal model. While recent advances in robust RL focus on Markov decision processes (MDPs), robust non-Markovian RL is limited to planning problem where the transitions in the uncertainty set are known. In this paper, we study the learning problem of robust offline non-Markovian RL. Specifically, when the nominal model admits a low-rank structure, we propose a new algorithm, featuring a novel dataset distillation and a lower confidence bound (LCB) design for robust values under different types of the uncertainty set. We also derive new dual forms for these robust values in non-Markovian RL, making our algorithm more amenable to practical implementation. By further introducing a novel type-I concentrability coefficient tailored for offline low-rank non-Markovian decision processes, we prove that our algorithm can find an $ε$-optimal robust policy using $O(1/ε^2)$ offline samples. Moreover, we extend our algorithm to the case when the nominal model does not have specific structure. With a new type-II concentrability coefficient, the extended algorithm also enjoys polynomial sample efficiency under all different types of the uncertainty set.

强化学习离线学习鲁棒性非马尔可夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。