arXiv:2504.11874cs.LG2025-04被引 2

提出可干预风险偏好的多智能体投资组合优化方法

Factor-MCLS: Multi-agent learning system with reward factor matrix and multi-critic framework for dynamic portfolio optimization

  • 引入奖励因子矩阵解析资产收益与风险
  • 多评论家框架提升对因子的理解能力
  • 支持投资者按风险偏好动态调整策略

传统深度强化学习(DRL)代理在动态投资组合优化中,通过分析奖励函数输出值来学习影响投资组合收益与风险的因素,并在训练环境中调整投资组合权重。然而,该方法存在重大局限:投资者难以根据对各资产的不同风险厌恶程度干预训练过程。这一困难源于另一个问题:现有DRL代理仅从奖励函数输出学习,可能无法深入理解影响投资组合收益与风险的关键因素。因此,目标投资组合权重的确定完全依赖于DRL代理自身。为解决上述问题,本文提出一种奖励因子矩阵,用于揭示投资组合中每项资产的收益与风险特征。同时,提出一种名为Factor-MCLS的新颖学习系统,采用多评论家框架,促进奖励因子矩阵的学习。在此机制下,基于多评论家框架中的评论家网络,我们在策略函数的训练目标函数中引入风险约束项。该约束项使投资者可根据自身对各资产的风险厌恶水平,动态干预DRL代理的训练过程。

原文摘要 · Abstract (English)

Typical deep reinforcement learning (DRL) agents for dynamic portfolio optimization learn the factors influencing portfolio return and risk by analyzing the output values of the reward function while adjusting portfolio weights within the training environment. However, it faces a major limitation where it is difficult for investors to intervene in the training based on different levels of risk aversion towards each portfolio asset. This difficulty arises from another limitation: existing DRL agents may not develop a thorough understanding of the factors responsible for the portfolio return and risk by only learning from the output of the reward function. As a result, the strategy for determining the target portfolio weights is entirely dependent on the DRL agents themselves. To address these limitations, we propose a reward factor matrix for elucidating the return and risk of each asset in the portfolio. Additionally, we propose a novel learning system named Factor-MCLS using a multi-critic framework that facilitates learning of the reward factor matrix. In this way, our DRL-based learning system can effectively learn the factors influencing portfolio return and risk. Moreover, based on the critic networks within the multi-critic framework, we develop a risk constraint term in the training objective function of the policy function. This risk constraint term allows investors to intervene in the training of the DRL agent according to their individual levels of risk aversion towards the portfolio assets.

强化学习投资组合多智能体风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。