arXiv:2410.04217cs.AIq-fin.PM2024-10被引 1

用自适应强化学习优化投资组合,动态调整资产配置更有效。

Improving Portfolio Optimization Results with Bandit Networks

  • 设计自适应折扣汤普森采样算法,应对奖励分布变化。
  • 在真实股市数据上,新方法使夏普比率提升20%以上。
  • 提出带状网络架构,解决组合优化计算瓶颈,适合量化交易者。

在强化学习中,多臂老虎机(MAB)问题已广泛应用于推荐系统、医疗和金融等领域。传统MAB算法通常假设奖励分布稳定,限制了其在非平稳现实环境中的应用。本文提出并评估了适用于非平稳环境的新颖老虎机算法。首先引入自适应折扣汤普森采样(ADTS)算法,通过放松折扣率和滑动窗口机制增强对奖励分布变化的响应能力。随后将该方法拓展至投资组合优化,提出组合型自适应折扣汤普森采样(CADTS)算法,解决了组合老虎机中的计算挑战,提升了动态资产配置效果。此外,提出新型架构——带状网络(Bandit Networks),整合ADTS与CADTS输出,缓解股票选择中的计算局限。基于真实金融市场数据的大量实验表明,这些算法与架构能有效适应动态环境并优化决策。例如,所提带状网络实例相较经典方法(如CAPM、等权、风险平价、马科维茨)表现更优,最佳网络的样本外夏普比率高出最优传统模型20%。

原文摘要 · Abstract (English)

In Reinforcement Learning (RL), multi-armed Bandit (MAB) problems have found applications across diverse domains such as recommender systems, healthcare, and finance. Traditional MAB algorithms typically assume stationary reward distributions, which limits their effectiveness in real-world scenarios characterized by non-stationary dynamics. This paper addresses this limitation by introducing and evaluating novel Bandit algorithms designed for non-stationary environments. First, we present the Adaptive Discounted Thompson Sampling (ADTS) algorithm, which enhances adaptability through relaxed discounting and sliding window mechanisms to better respond to changes in reward distributions. We then extend this approach to the Portfolio Optimization problem by introducing the Combinatorial Adaptive Discounted Thompson Sampling (CADTS) algorithm, which addresses computational challenges within Combinatorial Bandits and improves dynamic asset allocation. Additionally, we propose a novel architecture called Bandit Networks, which integrates the outputs of ADTS and CADTS, thereby mitigating computational limitations in stock selection. Through extensive experiments using real financial market data, we demonstrate the potential of these algorithms and architectures in adapting to dynamic environments and optimizing decision-making processes. For instance, the proposed bandit network instances present superior performance when compared to classic portfolio optimization approaches, such as capital asset pricing model, equal weights, risk parity, and Markovitz, with the best network presenting an out-of-sample Sharpe Ratio 20% higher than the best performing classical model.

投资组合强化学习动态优化带状网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。