用强化学习直接学投资策略,不估市场参数,效果优于传统方法。
Mean--Variance Portfolio Selection by Continuous-Time Reinforcement Learning: Algorithms, Regret Analysis, and Empirical Study
- 直接从数据学投资策略,跳过估计市场参数的步骤。
- 理论证明策略表现有保障,后悔值随时间呈亚线性下降。
- 实测在熊市中表现突出,显著优于基于模型的传统方法。
我们研究了在股票价格由可观测因素驱动的连续时间均值-方差投资组合选择问题,其中这些因素的过程系数未知。基于近期发展的扩散过程强化学习理论,提出一种通用的数据驱动强化学习方法,直接学习预承诺投资策略,无需学习或估计市场系数。针对无因子的多资产Black--Scholes市场,进一步设计了一种算法,并通过推导以夏普比率表示的亚线性后悔界,证明其性能保证。随后在标普500成分股上进行了广泛实证研究,将该算法与大量常用投资组合策略在多种常见指标下进行对比。结果表明,所提出的连续时间强化学习策略始终表现优异,尤其在波动剧烈的熊市中;且显著优于基于模型的连续时间方法。
原文摘要 · Abstract (English)
We study continuous-time mean--variance portfolio selection in markets where stock prices are diffusion processes driven by observable factors that are also diffusion processes, yet the coefficients of these processes are unknown. Based on the recently developed reinforcement learning (RL) theory for diffusion processes, we present a general data-driven RL approach that learns the pre-committed investment strategy directly without attempting to learn or estimate the market coefficients. For multi-stock Black--Scholes markets without factors, we further devise an algorithm and prove its performance guarantee by deriving a sublinear regret bound in terms of the Sharpe ratio. We then carry out an extensive empirical study implementing this algorithm to compare its performance and trading characteristics, evaluated under a host of common metrics, with a large number of widely employed portfolio allocation strategies on S\&P 500 constituents. The results demonstrate that the proposed continuous-time RL strategy is consistently among the best, especially in a volatile bear market, and decisively outperforms the model-based continuous-time counterparts by significant margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。