用强化学习训练的LLM模型,通过经济逻辑判断因子有效性
Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
- 用强化学习训练80亿参数模型,结合新闻和因子逻辑推理
- 在多资产池测试中超越基准策略,缓解因子衰减问题
- 适合量化投资研究者,尤其关注市场变化下的因子筛选
信号衰减和市场状态突变给非平稳市场中的数据驱动投资策略带来持续挑战。传统的时间序列与机器学习方法主要依赖历史相关性,当经济环境变化时往往难以泛化。尽管大语言模型(LLMs)具备处理非结构化信息的强大能力,但其通过明确经济推理支持量化因子筛选的潜力尚未被充分探索。现有基于因子的方法通常将阿尔法简化为数值时间序列,忽略了决定因子经济相关性的语义逻辑。本文提出Alpha-R1,一个通过强化学习训练的80亿参数推理模型,能够基于上下文感知的因子逻辑与实时新闻,评估因子在变化市场条件下的相关性,并根据上下文一致性选择性激活或禁用因子。在多个资产池的实证结果表明,Alpha-R1持续优于基准策略,并展现出对因子衰减更强的鲁棒性。完整实现与资源见 https://github.com/FinStep-AI/Alpha-R1。
原文摘要 · Abstract (English)
Signal decay and regime shifts pose recurring challenges for data-driven investment strategies in non-stationary markets. Conventional time-series and machine learning approaches, which rely primarily on historical correlations, often struggle to generalize when the economic environment changes. While large language models (LLMs) offer strong capabilities for processing unstructured information, their potential to support quantitative factor screening through explicit economic reasoning remains underexplored. Existing factor-based methods typically reduce alphas to numerical time series, overlooking the semantic rationale that determines when a factor is economically relevant. We propose Alpha-R1, an 8B-parameter reasoning model trained via reinforcement learning for context-aware alpha screening. Alpha-R1 reasons over factor logic and real-time news to evaluate alpha relevance under changing market conditions, selectively activating or deactivating factors based on contextual consistency. Empirical results across multiple asset pools show that Alpha-R1 consistently outperforms benchmark strategies and exhibits improved robustness to alpha decay. The full implementation and resources are available at https://github.com/FinStep-AI/Alpha-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。