用深度强化学习优化投资组合,让60/40策略更抗跌、收益更高。
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
- 用PPO算法动态调仓,结合后悔值优化的奖励函数。
- 在20个独立训练中平均提升收益,最大回撤降低12.3%。
- 适合量化交易、资产配置研究者,对交易成本敏感场景实用。
本文提出一种基于代理的增强方法,利用近端策略优化(PPO)改进已有高绩效投资组合策略。目标是通过PPO与预言家代理驱动的动态再平衡,提升传统的60/40基准组合(60%股票,40%债券)。采用基于后悔值的夏普比率奖励函数,并设计交易成本调度器以缓解手续费摩擦和信号损耗。引入前瞻性奖励函数,并使用循环块自举法生成合成数据,提升策略泛化能力。评估聚焦于收益与最大回撤两项指标。鉴于金融市场高度随机性,每期训练20个独立代理,取平均表现对比基准。结果表明,该方法不仅有效增强现有策略,且优于多个基线模型。
原文摘要 · Abstract (English)
This paper introduces a novel agent-based approach for enhancing existing portfolio strategies using Proximal Policy Optimization (PPO). Rather than focusing solely on traditional portfolio construction, our approach aims to improve an already high-performing strategy through dynamic rebalancing driven by PPO and Oracle agents. Our target is to enhance the traditional 60/40 benchmark (60% stocks, 40% bonds) by employing the Regret-based Sharpe reward function. To address the impact of transaction fee frictions and prevent signal loss, we develop a transaction cost scheduler. We introduce a future-looking reward function and employ synthetic data training through a circular block bootstrap method to facilitate the learning of generalizable allocation strategies. We focus on two key evaluation measures: return and maximum drawdown. Given the high stochasticity of financial markets, we train 20 independent agents each period and evaluate their average performance against the benchmark. Our method not only enhances the performance of the existing portfolio strategy through strategic rebalancing but also demonstrates strong results compared to other baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。