通过奖励跳跃记忆提升公式化因子发现的稳定性和效果
AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

- 引入奖励跳跃记忆机制,仅在公式完成时更新状态
- 使用随机粒子模拟未来收益,指导动作选择并提升预测精度
- 适合量化投资中需长期稳定表现的因子挖掘场景
公式化因子发现是依赖池的符号搜索问题,反馈主要在完整表达式评估后才可获得。这种延迟反馈导致两个耦合难题:保留的因子池无法保存完整的评估历史,且中间构建动作的价值不确定,因其后果取决于最终完成的公式。本文提出AlphaRJM,通过奖励跳跃记忆(Reward-Jump Memory)解决此问题——该机制在符号构建期间保持不变,仅在终端评估事件时根据实际池奖励与评估结果更新;同时引入动作条件的随机微分方程(SDE)回报批判器,用随机粒子表示未来折扣收益。粒子通过均值和不确定性指导动作选择,并通过结合能量-距离匹配、均值校准与跳跃正则化的分布贝尔曼目标进行学习。实验表明,AlphaRJM在多个股票池、预测周期和随机种子下均实现强而稳定的性能提升;消融实验证实了持久评估历史、随机回报建模与分布监督的互补作用。
原文摘要 · Abstract (English)
Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。