arXiv:2511.18850cs.CL2025-11ACL被引 10

用大模型驱动代码演化,自动发现可解释的金融预测信号。

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

  • 将因子表示为代码,通过大模型进行多轮推理与演化优化。
  • 在5个股票数据集上,预测准确率优于现有方法,泛化能力更强。
  • 适合对量化金融、可解释模型感兴趣的读者。

从高维且信噪比极低的金融数据中挖掘有效的预测信号(即“阿尔法”)仍是开放难题。尽管深度学习、遗传编程及近期的大语言模型(LLM)因子生成取得进展,现有方法仍局限于庞大的阿尔法搜索空间的狭窄区域。神经模型常产生难以解释且脆弱的模式,而符号或公式方法则往往得出冗余或经济上无依据的表达式,泛化性能差。这些范式虽形式不同,但共同局限在于无法实现广域、结构化且类人般的探索,兼顾逻辑一致性与创造性突破。为此,我们提出认知阿尔法挖掘框架(CogAlpha),结合代码级阿尔法表示、大模型驱动推理与进化搜索。将大模型视为自适应认知代理,通过多阶段提示与金融反馈,迭代地精炼、变异和重组阿尔法候选。该协同设计实现了更深层次思考、更丰富的结构多样性,以及经济可解释的阿尔法发现,并显著拓展有效搜索空间。在来自3个股票市场的5个股票数据集上的实验表明,CogAlpha持续发现具有更高预测准确性、鲁棒性与泛化能力的阿尔法,超越现有方法。结果凸显了将进化优化与基于大模型的推理相结合,在自动化与可解释阿尔法发现中的潜力。

原文摘要 · Abstract (English)

Discovering effective predictive signals, or "alphas," from financial data with high dimensionality and extremely low signal-to-noise ratio remains a difficult open problem. Despite progress in deep learning, genetic programming, and, more recently, large language model (LLM)-based factor generation, existing approaches still explore only a narrow region of the vast alpha search space. Neural models tend to produce opaque and fragile patterns, while symbolic or formula-based methods often yield redundant or economically ungrounded expressions that generalize poorly. Although different in form, these paradigms share a key limitation: none can conduct broad, structured, and human-like exploration that balances logical consistency with creative leaps. To address this gap, we introduce the Cognitive Alpha Mining Framework (CogAlpha), which combines code-level alpha representation with LLM-driven reasoning and evolutionary search. Treating LLMs as adaptive cognitive agents, our framework iteratively refines, mutates, and recombines alpha candidates through multi-stage prompts and financial feedback. This synergistic design enables deeper thinking, richer structural diversity, and economically interpretable alpha discovery, while greatly expanding the effective search space. Experiments on 5 stock datasets from 3 stock markets demonstrate that CogAlpha consistently discovers alphas with superior predictive accuracy, robustness, and generalization over existing methods. Our results highlight the promise of aligning evolutionary optimization with LLM-based reasoning for automated and explainable alpha discovery.

量化金融大模型进化算法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。