arXiv:2507.20263cs.LGcs.AI2025-07被引 6

通过专家公式相似度,让强化学习更高效挖掘量化投资因子。

Learning from Expert Factors: Trajectory-level Reward Shaping for Formulaic Alpha Mining

  • 用部分生成表达式与专家公式对比,提供中间奖励
  • 提升因子预测能力,排名信息系数提高9.29%
  • 计算效率大幅优化,时间复杂度从线性降为常数

强化学习已成功用于自动化挖掘可解释且盈利的公式化阿尔法因子,以构建投资策略。然而,现有方法受限于马尔可夫决策过程中的稀疏奖励,导致探索庞大符号搜索空间效率低下,并使训练过程不稳定。为此,提出一种轨迹级奖励塑造(TLRS)新方法:通过测量部分生成表达式与一组专家设计公式之间的子序列相似度,提供密集的中间奖励;同时引入奖励中心化机制以降低训练方差。在六个主要中、美股票指数上的大量实验表明,TLRS显著提升了挖掘因子的预测能力,相较现有基于潜力的塑造算法,排名信息系数提升9.29%。尤为关键的是,TLRS将相对于特征维度的时间复杂度从线性降至常数,较基于距离的基线实现重大计算效率提升。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has successfully automated the complex process of mining formulaic alpha factors, for creating interpretable and profitable investment strategies. However, existing methods are hampered by the sparse rewards given the underlying Markov Decision Process. This inefficiency limits the exploration of the vast symbolic search space and destabilizes the training process. To address this, Trajectory-level Reward Shaping (TLRS), a novel reward shaping method, is proposed. TLRS provides dense, intermediate rewards by measuring the subsequence-level similarity between partially generated expressions and a set of expert-designed formulas. Furthermore, a reward centering mechanism is introduced to reduce training variance. Extensive experiments on six major Chinese and U.S. stock indices show that TLRS significantly improves the predictive power of mined factors, boosting the Rank Information Coefficient by 9.29% over existing potential-based shaping algorithms. Notably, TLRS achieves a major leap in computational efficiency by reducing its time complexity with respect to the feature dimension from linear to constant, which is a significant improvement over distance-based baselines.

强化学习量化投资因子挖掘奖励塑造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。