提出新框架解决非指数折扣下的强化学习失效问题
Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting
- 用庞特里亚金原理替代贝尔曼递归,避免折扣失效
- 在多维双曲与生存折扣任务中显著提升稳定性和精度
- 适合研究人类偏好或长期生存决策的学者
大多数基于价值和演员-评论家的强化学习方法依赖贝尔曼递归,但在人类偏好和生存过程常见的非指数折扣下会失效。我们发现其崩溃具有结构性:指数折扣位于乘法性与时间齐性的脆弱交点,违反任一性质都会破坏标准动态规划。为此,我们提出庞特里亚金引导的直接策略优化(PG-DPO),一种变分框架,放弃递归,通过伴随-MC投影将庞特里亚金最大值原理与蒙特卡洛回溯结合,强制逐点哈密顿量最大化。在多维双曲折扣和生存折扣基准测试中,PG-DPO在方程驱动求解器和评论家基线发散时仍保持高精度与稳定性。
原文摘要 · Abstract (English)
Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential discounting common in human preferences and survival processes. We show the breakdown is structural: exponential discounting sits at a fragile intersection of multiplicativity and time homogeneity, and violating either property breaks standard dynamic programming. To overcome this, we propose Pontryagin-Guided Direct Policy Optimization (PG-DPO), a variational framework that abandons recursion and couples the Pontryagin Maximum Principle with Monte Carlo rollouts via an Adjoint-MC projection enforcing pointwise Hamiltonian maximization. Across multi-dimensional hyperbolic and survival-discount benchmarks, PG-DPO improves accuracy and stability where equation-driven solvers and critic-based baselines diverge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。