arXiv:2505.04553q-fin.MFcs.AI2025-05被引 3

用凸评分函数统一建模风险,实现更稳健的强化学习决策。

Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions

  • 通过扩展状态空间和引入辅助变量,解决风险优化的时间不一致性问题。
  • 算法在金融套利模拟中表现优异,有效控制风险并提升收益稳定性。
  • 理论无需马尔可夫过程连续性,适合复杂现实场景,如金融交易。

我们提出一种基于凸评分函数的风险敏感强化学习框架,涵盖方差、预期短缺、熵值风险度量及均值-风险效用等多种常见风险度量。为解决时间不一致性问题,引入扩展状态空间和辅助变量,将原问题重构为双状态优化。设计定制化演员-评论家算法,并建立理论近似保证。关键理论贡献在于:结果不依赖于马尔可夫决策过程的连续性。此外,提出受交替最小化启发的辅助变量采样方法,在特定条件下具有收敛性。通过金融领域的统计套利交易模拟实验验证方法有效性,表明算法在风险控制与收益提升方面表现良好。

原文摘要 · Abstract (English)

We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

强化学习风险控制金融应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。