用经济假设验证因子,让机器发现的金融因子既有效又可信。
FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation

- 以经济假设为起点,分三步验证因子是否真实反映预期条件。
- 在沪深300和标普500上表现优于现有方法,且跨市场周期稳定。
- 适合追求可解释性与稳健性的量化研究者和金融机构。
传统金融依赖专家基于经济逻辑手工构建因子。近期基于大模型的多智能体系统虽自动化了因子挖掘,但直接优化收益,极少检验生成因子是否仍符合其背后的经济假设。我们识别出数学形式与经济意义之间的不一致是收益导向自动化的结构性缺陷,导致因子模糊真实信号与虚假相关,且在市场结构变化时失效。为此,我们提出FaVOR(通过可观测推理进行因子验证),一种围绕假设层面证据重构的智能体框架。不同于传统的假设到公式的跳跃,FaVOR通过三阶段一致性循环,将数学形式始终与经济逻辑绑定:(1) 分解将宽泛经济假设拆分为独立可观测条件;(2) 验证检查每个因子是否反映其对应条件;(3) 整合生成结构可解释的复合因子。在2025年沪深300和标普500数据上,FaVOR超越现有基线,且跨市场周期保持有效性。结果表明,基于假设的因子发现能产生可解释、抗周期、经济忠实的信号。代码已开源:https://github.com/damilab/FaVOR。
原文摘要 · Abstract (English)
Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort. However, these automated approaches optimize directly for returns and rarely check whether a generated factor still expresses the economic hypothesis that motivated it. We identify this inconsistency between mathematical form and economic meaning as a structural failure mode of return-oriented automation. The resulting factors blur the line between real signals and spurious correlations and break down across regime shifts. We propose FaVOR (Factor Validation through Observable Reasoning), an agentic framework that restructures factor mining around hypothesis-level evidence rather than return outcomes. In place of the standard hypothesis-to-formula leap, FaVOR enforces a three-stage consistency loop tying mathematical form to economic rationale throughout. (1) Decomposition splits a broad economic hypothesis into independent observable conditions. (2) Validation checks whether each factor reflects its intended condition. (3) Integration merges them into a composite whose structure remains interpretable. On the CSI 500 and S&P 500 in 2025, FaVOR outperforms existing baselines while remaining effective across regimes. FaVOR shows that hypothesis-grounded factor discovery produces signals that are interpretable by construction, regime-robust, and economically faithful. The code is available at https://github.com/damilab/FaVOR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。