提出新算法DROP,用乐观悲观机制提升强化学习性能。
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
- 从控制即推理推导出带正负偏差的更新规则
- 在动态任务中表现优于启发式模型,接近顶尖算法
- 适合研究生物启发强化学习或鲁棒性优化的学者
强化学习中,时序差分(TD)误差与多巴胺神经元放电率相关。研究发现,多巴胺神经元对TD误差的响应具有乐观或悲观特性,这可解释为一种分布式强化学习。现有启发式模型对正负TD误差采用不对称学习率,但缺乏理论基础,且未知是否具备实际效果。本文提出一种基于控制即推理推导的理论严谨模型,通过集成学习引入规律性乐观与悲观,估计分布值函数作为评论家。基于其均值,策略在演员中被优化。该算法称为DROP(分布与规律性乐观悲观)。在动态任务上对比实验表明,尽管启发式模型表现不佳,DROP在所有任务中均表现出色,具有高泛化性,性能媲美当前最优算法。结果表明,DROP是一种能激发乐观悲观潜力的新模型。
原文摘要 · Abstract (English)
In reinforcement learning (RL), temporal difference (TD) error is known to be related to the firing rate of dopamine neurons. It has been observed that each dopamine neuron does not behave uniformly, but each responds to the TD error in an optimistic or pessimistic manner, interpreted as a kind of distributional RL. To explain such a biological data, a heuristic model has also been introduced with learning rates asymmetric for the positive and negative TD errors. However, this heuristic model is not theoretically-grounded and unknown whether it can work as a RL algorithm. This paper therefore introduces a novel theoretically-grounded model with optimism and pessimism, which is derived from control as inference. In combination with ensemble learning, a distributional value function as a critic is estimated from regularly introduced optimism and pessimism. Based on its central value, a policy in an actor is improved. This proposed algorithm, so-called DROP (distributional and regular optimism and pessimism), is compared on dynamic tasks. Although the heuristic model showed poor learning performance, DROP demonstrated excellent performance in all tasks with high generality. In addition, DROP achieved learning performance comparable to the state-of-the-art algorithms. In other words, it was suggested that DROP is a new model that can elicit the potential contributions of optimism and pessimism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。