首次为时间不一致偏好设计可落地的强化学习算法
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
- 提出最优策略为一步非平稳结构,简化决策逻辑
- 设计首个无需环境模型的评估与Q-learning算法,保证收敛
- 适合研究人类行为建模与非理性决策的学者
时间不一致偏好——即个体更倾向即时小奖励而非延迟大奖励——是人类和动物决策的核心特征。拟双曲(Quasi-Hyperbolic, QH)折扣模型能有效刻画此类行为,但其在强化学习框架中的应用受限。本文填补了预承诺代理在QH偏好下的理论与算法空白:首先,首次严格证明最优策略退化为简单的一步非平稳形式;其次,提出首个实用的无模型算法,用于策略评估与Q学习,并具备可证明的收敛性。研究成果为将QH偏好融入强化学习提供了基础支撑。
原文摘要 · Abstract (English)
Time-inconsistent preferences, where agents favor smaller-sooner over larger-later rewards, are a key feature of human and animal decision-making. Quasi-Hyperbolic (QH) discounting provides a simple yet powerful model for this behavior, but its integration into the reinforcement learning (RL) framework has been limited. This paper addresses key theoretical and algorithmic gaps for precommitted agents with QH preferences. We make two primary contributions: (i) we formally characterize the structure of the optimal policy, proving for the first time that it reduces to a simple one-step non-stationary form; and (ii) we design the first practical, model-free algorithms for both policy evaluation and Q-learning in this setting, both with provable convergence guarantees. Our results provide foundational insights for incorporating QH preferences in RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。