无先验的黑盒非平稳强化学习难奏效,检测机制常失效。
Is Prior-Free Black-Box Non-Stationary Reinforcement Learning Feasible?
- 用随机重启作为基准,验证非平稳性检测失效
- MASTER算法在合理时长下始终无法触发检测,表现如随机重启
- 基于快速变化检测的方法更鲁棒,显著优于MASTER
我们研究无先验知识下的非平稳强化学习(NS-RL)可行性。以先进黑盒算法MASTER为例,分析其能否实现宣称目标。结果表明,对于实际选择的时长,MASTER的非平稳性检测机制无法被触发,性能等同于随机重启算法。此外,尽管其遗憾界为阶最优,但在不合理大的时长前仍高于最坏情况线性遗憾。通过在分段平稳多臂赌博机场景中测试MASTER,并与随机重启及快速变化检测重启方法对比,验证了上述结论。提出一个已知非平稳性的简单、阶最优随机重启算法作为基线。模拟显示,采用快速变化检测的方法更具鲁棒性,持续优于MASTER及其他随机重启方法。
原文摘要 · Abstract (English)
We study the problem of Non-Stationary Reinforcement Learning (NS-RL) without prior knowledge about the system's non-stationarity. A state-of-the-art, black-box algorithm, known as MASTER, is considered, with a focus on identifying the conditions under which it can achieve its stated goals. Specifically, we prove that MASTER's non-stationarity detection mechanism is not triggered for practical choices of horizon, leading to performance akin to a random restarting algorithm. Moreover, we show that the regret bound for MASTER, while being order optimal, stays above the worst-case linear regret until unreasonably large values of the horizon. To validate these observations, MASTER is tested for the special case of piecewise stationary multi-armed bandits, along with methods that employ random restarting, and others that use quickest change detection to restart. A simple, order optimal random restarting algorithm, that has prior knowledge of the non-stationarity is proposed as a baseline. The behavior of the MASTER algorithm is validated in simulations, and it is shown that methods employing quickest change detection are more robust and consistently outperform MASTER and other random restarting approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。