arXiv:2412.05906stat.MLcs.LG2024-12

用强化学习求解离散LQ控制,发现最优策略是高斯型。

Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application

  • 以熵衡量探索代价,推导出最优反馈策略为高斯型
  • 将结果应用于资产-负债管理,证明算法可提升策略并收敛
  • 通过数值模拟验证理论,适合金融控制方向研究者

我们采用强化学习(RL)研究离散时间线性二次(LQ)控制模型。通过熵度量探索成本,证明该问题的最优反馈策略必为高斯型。随后,将离散时间LQ模型结果应用于离散时间均值-方差资产-负债管理问题,证明所提RL算法具有策略改进性和收敛性。最后,通过数值例子结合仿真验证了前述理论结果的有效性。

原文摘要 · Abstract (English)

We study the discrete-time linear-quadratic (LQ) control model using reinforcement learning (RL). Using entropy to measure the cost of exploration, we prove that the optimal feedback policy for the problem must be Gaussian type. Then, we apply the results of the discrete-time LQ model to solve the discrete-time mean-variance asset-liability management problem and prove our RL algorithm's policy improvement and convergence. Finally, a numerical example sheds light on the theoretical results established using simulations.

强化学习LQ控制金融建模策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。