arXiv:2409.05144q-fin.CPcs.AI2024-09被引 24

用改进的REINFORCE算法挖掘稳定可解释的量化投资因子。

QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE

  • 基于蒙特卡洛采样改进的REINFORCE算法,降低策略梯度方差。
  • 在真实资产数据上提升收益相关性3.83%,超额收益能力更强。
  • 适合追求可解释性与稳定性的量化投资研究者使用。

Alpha因子挖掘旨在从历史金融数据中发现可用于预测资产收益并获取超额利润的投资信号。尽管深度学习方法强大,但缺乏可解释性,在风险敏感的真实市场中难以被接受。公式化因子因具备可解释性而更受青睐,但搜索空间复杂,亟需高效探索方法。近期一种基于深度强化学习生成公式化因子的框架受到学界与业界广泛关注。本文指出,原采用的近端策略优化(PPO)在因子挖掘中存在若干问题。为此提出一种基于经典REINFORCE算法的新方法:利用蒙特卡洛采样估计策略梯度,虽具无偏但高方差特性;然而,底层状态转移函数遵循狄拉克分布,环境变异性极低,有助于缓解方差问题,使REINFORCE更适用。设计专用基线以理论降低REINFORCE固有的高方差。此外,引入信息比率作为奖励塑形机制,鼓励生成适应市场波动变化的稳定因子。在真实资产数据上的评估表明,该方法将因子与收益的相关性提升3.83%,且获得超额收益的能力优于最新方法,结果与理论预期一致。

原文摘要 · Abstract (English)

Alpha factor mining aims to discover investment signals from the historical financial market data, which can be used to predict asset returns and gain excess profits. Powerful deep learning methods for alpha factor mining lack interpretability, making them unacceptable in the risk-sensitive real markets. Formulaic alpha factors are preferred for their interpretability, while the search space is complex and powerful explorative methods are urged. Recently, a promising framework is proposed for generating formulaic alpha factors using deep reinforcement learning, and quickly gained research focuses from both academia and industries. This paper first argues that the originally employed policy training method, i.e., Proximal Policy Optimization (PPO), faces several important issues in the context of alpha factors mining. Herein, a novel reinforcement learning algorithm based on the well-known REINFORCE algorithm is proposed. REINFORCE employs Monte Carlo sampling to estimate the policy gradient-yielding unbiased but high variance estimates. The minimal environmental variability inherent in the underlying state transition function, which adheres to the Dirac distribution, can help alleviate this high variance issue, making REINFORCE algorithm more appropriate than PPO. A new dedicated baseline is designed to theoretically reduce the commonly suffered high variance of REINFORCE. Moreover, the information ratio is introduced as a reward shaping mechanism to encourage the generation of steady alpha factors that can better adapt to changes in market volatility. Evaluations on real assets data indicate the proposed algorithm boosts correlation with returns by 3.83\%, and a stronger ability to obtain excess returns compared to the latest alpha factors mining methods, which meets the theoretical results well.

量化投资强化学习因子挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。