arXiv:2409.10096cs.LGq-fin.CP2024-09被引 5

提出动态风险度量框架,提升强化学习在不确定环境下的鲁棒性决策能力。

Robust Reinforcement Learning with Dynamic Distortion Risk Measures

  • 用Wasserstein球内所有模型构建环境不确定性,结合动态扭曲风险度量
  • 通过神经网络估计风险度量,基于分位数表示推导策略梯度公式
  • 适用于金融投资等需考虑风险与不确定性的场景,尤其适合对鲁棒性要求高的应用

在强化学习环境中,智能体的最优策略高度依赖其风险偏好和训练环境的模型动态。这两者共同影响智能体在测试环境中的决策质量与时序一致性。本文提出一种解决鲁棒风险感知强化学习问题的框架,同时考虑环境不确定性与风险,采用一类动态鲁棒扭曲风险度量。通过在参考模型的Wasserstein球内考虑所有可能模型引入鲁棒性。利用严格一致评分函数,通过神经网络估计此类动态鲁棒风险度量,并基于扭曲风险度量的分位数表示推导策略梯度公式,构建了相应的演员-评论家算法来求解该类问题。在投资组合配置示例中验证了所提算法的有效性。

原文摘要 · Abstract (English)

In a reinforcement learning (RL) setting, the agent's optimal strategy heavily depends on her risk preferences and the underlying model dynamics of the training environment. These two aspects influence the agent's ability to make well-informed and time-consistent decisions when facing testing environments. In this work, we devise a framework to solve robust risk-aware RL problems where we simultaneously account for environmental uncertainty and risk with a class of dynamic robust distortion risk measures. Robustness is introduced by considering all models within a Wasserstein ball around a reference model. We estimate such dynamic robust risk measures using neural networks by making use of strictly consistent scoring functions, derive policy gradient formulae using the quantile representation of distortion risk measures, and construct an actor-critic algorithm to solve this class of robust risk-aware RL problems. We demonstrate the performance of our algorithm on a portfolio allocation example.

强化学习风险感知鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。