arXiv:2606.11798q-fin.CPcs.LG2026-06被引 2

提出新算法统一求解时间不一致控制问题的均衡策略。

Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems

论文配图:Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems
图 1 · 摘自论文原文
  • 分两阶段迭代:先学最优策略,再更新辅助函数
  • 在均值-方差与非指数贴现追踪组合中验证有效
  • 适用于金融领域多种时间不一致场景

本文提出一种连续时间、模型无关的强化学习算法,用于求解一般时间不一致控制问题中的确定性均衡策略。通过扩展的哈密顿-雅可比-贝尔曼系统,将原问题转化为等价的两阶段问题。第一阶段,在给定辅助函数下,采用确定性策略梯度方法求解一个辅助的时间一致控制问题的最优策略;第二阶段,在更新策略后,利用内层不动点迭代和鞅表征学习辅助函数。理论方面,在较弱模型假设下证明了内层不动点迭代的收敛性。通过跨两阶段的演员-评论家式迭代,该算法旨在统一求解不同来源的时间不一致性下的均衡策略。在两类经典金融应用中验证了算法优越性:均值-方差投资组合管理与非指数贴现下的最优跟踪组合。

原文摘要 · Abstract (English)

In this paper, we develop a continuous-time model-free reinforcement learning algorithm to learn deterministic equilibrium policies in general time-inconsistent control problems. Utilizing the extended Hamilton-Jacobi-Bellman system, we recast the original time-inconsistent problem into an equivalent two-stage problem. In the first stage, for given auxiliary functions, we employ the deterministic policy gradient approach to learn an optimal policy in an auxiliary time-consistent control problem. In the second stage, given the updated policy, we exploit the inner fixed point iterations and some martingale characterizations to learn the auxiliary functions. As a theoretical contribution, we provide some mild model assumptions and establish the convergence of inner fixed point iterations. By repeating this actor-critic style of iterations across two stages, our algorithm aims to learn the equilibrium under different sources of time-inconsistency in a unified manner. The superior effectiveness of the proposed algorithm are illustrated in two classical financial applications with time-inconsistency: mean-variance portfolio management and optimal tracking portfolio under non-exponential discounting.

强化学习时间不一致金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。