用结构化语义空间提升大模型选股因子挖掘效率与可控性
AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining

- 构建事件-上下文-属性-方向-输出五元组语义空间,规范因子生成逻辑
- 在沪深股市实验中实现显著预测力与组合收益,超越传统方法
- 支持多模型适配,同一语义方案下不同大模型表现一致,鲁棒性强
自动化选股因子挖掘正越来越多采用大语言模型(LLM)代理进行因子生成与迭代发现。然而现有系统通常将因子构造与搜索决策全权交由代理,缺乏显式的探索空间或系统的导航机制,导致探索过程隐式且难以控制与优化。本文提出AlphaSchema,构建并探索用于选股因子挖掘的结构化交易语义空间。该空间中每个点为包含事件、上下文、属性、方向和输出的语义计划(schema plan),在实现前即定义候选因子的语义。AlphaSchema解耦探索与实现:一个LLM将选定的语义计划转化为可执行因子,评估回报被累积以学习语义空间上的代理模型。迭代选择机制利用该模型平衡全局探索、代理引导的利用与局部变异。在中国股票市场上的实验表明,AlphaSchema发现的因子池具有优异的预测能力和投资组合表现。进一步分析显示,语义搜索过程能有效覆盖多样区域,并逐步聚焦高回报区域;同一语义计划由不同LLM实现时表现出相近的预测质量,说明在本框架下因子挖掘质量对具体大模型选择具有较强鲁棒性。
原文摘要 · Abstract (English)
Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space. As a result, exploration remains largely implicit and difficult to control or optimize systematically. We introduce AlphaSchema, which constructs and explores a structured space of trading semantics for alpha mining. Each point in this space is a schema plan composed of Event, Context, Qualities, Direction, and Output, specifying the semantics of a candidate factor before implementation. AlphaSchema decouples exploration from implementation: an LLM translates selected schema plans into executable factors, while evaluated rewards are accumulated to learn a surrogate model over the semantic space. An iterative selection mechanism uses this model to balance global exploration, surrogate-guided exploitation, and local mutation. Experiments on the Chinese stock market show that AlphaSchema discovers factor pools with strong predictive and portfolio performance. Further analyses show that the semantic search process navigates diverse regions while increasingly allocating evaluations toward high-reward regions, and that implementations of the same schema plans by different LLMs exhibit comparable predictive quality, suggesting that alpha mining quality is largely robust to the choice of LLM within our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。