针对非马尔可夫奖励的强化学习,提出首个结构化高效探索算法。
Provably Efficient Exploration in Reward Machines with Low Regret
- 利用奖励机结构设计模型化强化学习算法
- 理论证明其后悔值优于传统无结构方法
- 适合有高层任务知识的智能体高效学习
我们研究具有非马尔可夫奖励的决策过程中的强化学习(RL),其中任务的高层知识以奖励机形式提供给学习者。考虑动态未知的概率性奖励机,在平均奖励准则下评估学习性能,使用后悔(regret)衡量。主要贡献是一种基于模型的强化学习算法,能有效利用奖励机带来的结构信息。我们推导出该算法的高概率、非渐近后悔上界,并展示其在后悔表现上相比未利用结构的现有算法有显著提升。同时给出了该设置下的后悔下界。据我们所知,该算法是首个为概率性奖励机专门设计并分析后悔的尝试。
原文摘要 · Abstract (English)
We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge of the task in the form of reward machines is available to the learner. We consider probabilistic reward machines with initially unknown dynamics, and investigate RL under the average-reward criterion, where the learning performance is assessed through the notion of regret. Our main algorithmic contribution is a model-based RL algorithm for decision processes involving probabilistic reward machines that is capable of exploiting the structure induced by such machines. We further derive high-probability and non-asymptotic bounds on its regret and demonstrate the gain in terms of regret over existing algorithms that could be applied, but obliviously to the structure. We also present a regret lower bound for the studied setting. To the best of our knowledge, the proposed algorithm constitutes the first attempt to tailor and analyze regret specifically for RL with probabilistic reward machines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。