arXiv:2411.18612cs.LGcs.AI2024-11ICML被引 8

提出结构化正则化框架,提升离线强化学习在动态变化下的鲁棒性。

Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization

  • 引入线性结构的f散度正则化,约束转移概率的潜在关系。
  • 理论证明算法近似最优,且性能依赖数据对最优策略路径的覆盖程度。
  • 适合需要应对环境动态扰动的离线强化学习场景。

为提升策略在动态变化下的鲁棒性,本文提出鲁棒正则化马尔可夫决策过程(RRMDP),通过在价值函数中对转移动态施加正则化实现。现有方法多采用无结构正则化,可能导致在不现实转移下产生保守策略。为此,本文提出d-矩形线性RRMDP(d-RRMDP)框架,将潜在结构同时引入转移核与正则化项。聚焦离线强化学习场景,即智能体从预收集的数据集中学习策略。提出鲁棒正则化悲观值迭代(R2PVI)算法,结合线性函数逼近,在基于f散度的转移核正则化下实现鲁棒策略学习。提供实例相关的次优间隙上界,表明其性能取决于数据集对最优鲁棒策略在鲁棒可接受转移下访问的状态-动作空间的覆盖程度。建立信息论下界验证该算法接近最优。数值实验表明,R2PVI能学习到鲁棒策略,且计算效率优于基线方法。

原文摘要 · Abstract (English)

The Robust Regularized Markov Decision Process (RRMDP) is proposed to learn policies robust to dynamics shifts by adding regularization to the transition dynamics in the value function. Existing methods mostly use unstructured regularization, potentially leading to conservative policies under unrealistic transitions. To address this limitation, we propose a novel framework, the $d$-rectangular linear RRMDP ($d$-RRMDP), which introduces latent structures into both transition kernels and regularization. We focus on offline reinforcement learning, where an agent learns policies from a precollected dataset in the nominal environment. We develop the Robust Regularized Pessimistic Value Iteration (R2PVI) algorithm that employs linear function approximation for robust policy learning in $d$-RRMDPs with $f$-divergence based regularization terms on transition kernels. We provide instance-dependent upper bounds on the suboptimality gap of R2PVI policies, demonstrating that these bounds are influenced by how well the dataset covers state-action spaces visited by the optimal robust policy under robustly admissible transitions. We establish information-theoretic lower bounds to verify that our algorithm is near-optimal. Finally, numerical experiments validate that R2PVI learns robust policies and exhibits superior computational efficiency compared to baseline methods.

强化学习离线学习鲁棒性正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。