提出DARLING算法,提升非平稳强化学习的鲁棒性与性能。
DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees

- 通过检测变化点动态调整策略,无需提前知晓变化时机。
- 在表格和线性MDP中均达到最优动态遗憾下界,理论性能领先。
- 适用于多种非平稳环境,适合需要稳定决策的机器人、金融场景。
我们研究在无先验非平稳信息条件下,非平稳有限时域周期马尔可夫决策过程(MDPs)中的无模型强化学习(RL)。聚焦于分段平稳(PS)设置,其中奖励和转移动态可在未知时间发生改变。我们重新审视现有最先进方法,揭示其理论与实践局限,重塑性能保证的现状。为刻画问题难度,首次建立了表格型和线性MDP中PS-RL的极小极大下界。随后提出检测增强强化学习(DARLING),一种适用于表格型和线性MDP的模块化封装方法,无需知晓变化点。在表格型MDP中,在变化点可分离性和可达性条件下,DARLING改进了已知最佳动态遗憾界,并匹配我们的极小极大下界。在线性MDP中,当相关可达性参数已知时,DARLING匹配极小极大下界,分析揭示了该设定与表格情形间的结构性障碍。最后,通过在多样非平稳基准上的广泛实验,表明DARLING始终优于当前最先进方法。
原文摘要 · Abstract (English)
We study model-free reinforcement learning (RL) in non-stationary finite-horizon episodic Markov decision processes (MDPs) without prior knowledge of the non-stationarity. We focus on the piecewise stationary (PS) setting, where both rewards and transition dynamics can change at unknown times. We first revisit existing state-of-the-art approaches and identify theoretical and practical limitations that change the current landscape of performance guarantees. To characterize the difficulty of the problem, we establish the first minimax lower bounds for PS-RL in tabular and linear MDPs. We then introduce Detection Augmented Reinforcement Learning (DARLING), a modular wrapper for PS-RL that applies to both tabular and linear MDPs, without knowledge of the changes. In tabular MDPs, under change-point separability and reachability conditions, DARLING improves the best known dynamic regret bounds and matches our minimax lower bound. In linear MDPs, DARLING matches the minimax lower bound when the relevant reachability parameters are known, and our analysis clarifies the structural obstacles that distinguish this setting from the tabular case. Finally, through extensive experimentation across diverse non-stationary benchmarks, we show that DARLING consistently surpasses the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。