arXiv:2501.06832cs.LGcs.MA2025-01被引 19

用分层多智能体强化学习提升投资组合动态优化效果

A novel multi-agent dynamic portfolio optimization learning system based on hierarchical deep reinforcement learning

  • 设计分层架构,主智能体与辅助智能体协同探索高收益低波动策略
  • 在稀疏正回报环境下显著提升风险调整后收益,克服维度灾难
  • 适合关注智能投顾、量化交易系统优化的研究者和从业者

深度强化学习(DRL)被广泛用于解决投资组合优化问题。DRL智能体通过与环境的无监督交互获取知识并做出决策,无需掌握资产联合动态。目前最常用的DRL算法是结合演员-评论家框架与深度函数逼近器的方法。然而,我们发现使用该方法训练时,智能体的风险调整后盈利能力提升不明显。主要源于两个问题:正回报稀疏性与维度灾难。这些问题阻碍了智能体全面学习训练环境中的资产价格变化模式,导致无法有效探索动态投资组合优化策略以提升风险调整后收益。为此,本文提出一种新型多智能体分层深度强化学习(HDRL)算法框架。在此框架下,多个智能体作为学习系统协同工作。具体而言,通过设计一个与执行智能体协作的辅助智能体,共同聚焦于正回报且方差低的动作空间,探索更高风险调整收益的策略。该方法有效缓解了维度灾难问题,并提升了稀疏正回报环境下的训练效率。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) has been extensively used to address portfolio optimization problems. The DRL agents acquire knowledge and make decisions through unsupervised interactions with their environment without requiring explicit knowledge of the joint dynamics of portfolio assets. Among these DRL algorithms, the combination of actor-critic algorithms and deep function approximators is the most widely used DRL algorithm. Here, we find that training the DRL agent using the actor-critic algorithm and deep function approximators may lead to scenarios where the improvement in the DRL agent's risk-adjusted profitability is not significant. We propose that such situations primarily arise from the following two problems: sparsity in positive reward and the curse of dimensionality. These limitations prevent DRL agents from comprehensively learning asset price change patterns in the training environment. As a result, the DRL agents cannot explore the dynamic portfolio optimization policy to improve the risk-adjusted profitability in the training process. To address these problems, we propose a novel multi-agent Hierarchical Deep Reinforcement Learning (HDRL) algorithmic framework in this research. Under this framework, the agents work together as a learning system for portfolio optimization. Specifically, by designing an auxiliary agent that works together with the executive agent for optimal policy exploration, the learning system can focus on exploring the policy with higher risk-adjusted return in the action space with positive return and low variance. In this way, we can overcome the issue of the curse of dimensionality and improve the training efficiency in the positive reward sparse environment.

强化学习投资组合优化多智能体分层学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。