arXiv:2412.21004cs.LGcs.RO2024-12被引 1

将韦伯-费希纳定律引入强化学习,提升早期奖励获取与惩罚抑制能力。

Weber-Fechner Law in Temporal Difference learning derived from Control as Inference

  • 基于控制即推理框架,推导出非线性更新规则
  • 发现感知更新强度随价值函数增大而衰减的韦伯-费希纳规律
  • 实验验证算法加速奖励积累并持续抑制惩罚

本文研究一种基于时序差分(TD)误差的新型非线性更新规则。标准强化学习中,TD误差与更新幅度呈线性关系,对所有奖励无差别处理。然而,最新生物研究表明,TD误差与更新幅度之间存在非线性关系,导致策略偏向乐观或悲观。这种由非线性引起的偏差在生物学习中可能具有实用价值且被有意保留。为此,本文聚焦于控制即推理(control as inference)框架,该框架可统一多种强化学习与最优控制方法。特别地,分析了从控制即推理推导标准强化学习时需近似忽略的不可计算非线性项。结果发现韦伯-费希纳定律(WFL):对刺激变化(即TD误差)的感知(即更新程度)会随刺激强度(即价值函数)增加而减弱。为验证其在强化学习中的效用,提出一种基于奖惩框架的实用实现,并重构最优性定义。分析表明,该方法可实现两项优势:一是在早期快速提升奖励,二是在学习过程中充分抑制惩罚。通过仿真与机器人实验验证,所提带韦伯-费希纳定律的强化学习算法确实实现了加速奖励最大化启动并持续抑制惩罚的效果。

原文摘要 · Abstract (English)

This paper investigates a novel nonlinear update rule based on temporal difference (TD) errors in reinforcement learning (RL). The update rule in the standard RL states that the TD error is linearly proportional to the degree of updates, treating all rewards equally without no bias. On the other hand, the recent biological studies revealed that there are nonlinearities in the TD error and the degree of updates, biasing policies optimistic or pessimistic. Such biases in learning due to nonlinearities are expected to be useful and intentionally leftover features in biological learning. Therefore, this research explores a theoretical framework that can leverage the nonlinearity between the degree of the update and TD errors. To this end, we focus on a control as inference framework, since it is known as a generalized formulation encompassing various RL and optimal control methods. In particular, we investigate the uncomputable nonlinear term needed to be approximately excluded in the derivation of the standard RL from control as inference. By analyzing it, Weber-Fechner law (WFL) is found, namely, perception (a.k.a. the degree of updates) in response to stimulus change (a.k.a. TD error) is attenuated by increase in the stimulus intensity (a.k.a. the value function). To numerically reveal the utilities of WFL on RL, we then propose a practical implementation using a reward-punishment framework and modifying the definition of optimality. Analysis of this implementation reveals that two utilities can be expected i) to increase rewards to a certain level early, and ii) to sufficiently suppress punishment. We finally investigate and discuss the expected utilities through simulations and robot experiments. As a result, the proposed RL algorithm with WFL shows the expected utilities that accelerate the reward-maximizing startup and continue to suppress punishments during learning.

强化学习韦伯-费希纳非线性更新控制即推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。