arXiv:2409.10583cs.LGcs.AI2024-09被引 2

提出首个无需模型的马尔可夫完美均衡算法,解决人类延迟满足偏差问题。

Reinforcement Learning with Quasi-Hyperbolic Discounting

论文配图:Reinforcement Learning with Quasi-Hyperbolic Discounting
图 1 · 摘自论文原文
  • 基于双时间尺度分析设计无模型算法,求解马尔可夫完美均衡
  • 证明算法收敛时极限必为马尔可夫完美均衡,理论严谨可靠
  • 在随机需求库存系统上验证有效,适合有时间不一致行为的研究者

强化学习传统上采用指数折扣或平均奖励框架,主要因其数学可处理性。然而,这些框架难以准确刻画人类行为中对即时回报的偏好。拟双曲(QH)折扣是一种简单替代,能更好建模此偏好。与传统折扣不同,从时间 $t_1$ 和 $t_2$ 出发的最优策略可能不同,导致天真或急躁的未来自我偏离初始最优策略,造成整体回报下降。为防止此行为,可采用锚定在马尔可夫完美均衡(MPE)的策略。本文首次提出一种无模型算法以寻找 MPE。通过双时间尺度分析,证明若算法收敛,则极限必为 MPE。数值验证在具有随机需求的标准库存系统上完成,结果支持该理论结论。本工作显著推进了强化学习在实际场景中的应用。

原文摘要 · Abstract (English)

Reinforcement learning has traditionally been studied with exponential discounting or the average reward setup, mainly due to their mathematical tractability. However, such frameworks fall short of accurately capturing human behavior, which has a bias towards immediate gratification. Quasi-Hyperbolic (QH) discounting is a simple alternative for modeling this bias. Unlike in traditional discounting, though, the optimal QH-policy, starting from some time $t_1,$ can be different to the one starting from $t_2.$ Hence, the future self of an agent, if it is naive or impatient, can deviate from the policy that is optimal at the start, leading to sub-optimal overall returns. To prevent this behavior, an alternative is to work with a policy anchored in a Markov Perfect Equilibrium (MPE). In this work, we propose the first model-free algorithm for finding an MPE. Using a two-timescale analysis, we show that, if our algorithm converges, then the limit must be an MPE. We also validate this claim numerically for the standard inventory system with stochastic demands. Our work significantly advances the practical application of reinforcement learning.

强化学习马尔可夫均衡行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。