用连续时间强化学习直接学最优资产配置策略,不依赖环境建模。
Continuous-Time Reinforcement Learning for Asset-Liability Management
- 基于LQ框架的无模型强化学习算法,动态匹配资产与负债。
- 200个市场场景下平均收益超所有对比方法,初期增长快且持续领先。
- 无需复杂网络或参数估计,直接优化策略,适合金融风控场景。
本文提出一种新的资产-负债管理(ALM)方法,采用连续时间强化学习(RL)结合线性二次型(LQ)公式,同时考虑中期和终期目标。我们开发了一种无需模型、基于策略梯度的软演员-评论家算法,专用于动态协调资产与负债。为在最小调参条件下实现探索与利用的平衡,引入演员的自适应探索和评论家的调度式探索。实证研究将该方法与两种改进的传统金融策略、一种基于模型的连续时间RL方法以及三种前沿强化学习算法进行对比。在200个随机市场情景下评估,本方法获得更高平均回报,初始阶段快速提升并持续保持优势。其优异表现并非源于复杂神经网络或改进参数估计,而是通过直接学习最优ALM策略而实现,无需学习环境。
原文摘要 · Abstract (English)
This paper proposes a novel approach for Asset-Liability Management (ALM) by employing continuous-time Reinforcement Learning (RL) with a linear-quadratic (LQ) formulation that incorporates both interim and terminal objectives. We develop a model-free, policy gradient-based soft actor-critic algorithm tailored to ALM for dynamically synchronizing assets and liabilities. To ensure an effective balance between exploration and exploitation with minimal tuning, we introduce adaptive exploration for the actor and scheduled exploration for the critic. Our empirical study evaluates this approach against two enhanced traditional financial strategies, a model-based continuous-time RL method, and three state-of-the-art RL algorithms. Evaluated across 200 randomized market scenarios, our method achieves higher average rewards than all alternative strategies, with rapid initial gains and sustained superior performance. The outperformance stems not from complex neural networks or improved parameter estimation, but from directly learning the optimal ALM strategy without learning the environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。