arXiv:2602.18481q-fin.TRcs.AI2026-02KDD

用可复现的因子生成替代随机交易,让大模型真正像量化研究员一样思考。

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

论文配图:AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models
图 1 · 摘自论文原文
  • 让大模型生成可执行的金融因子,而非直接输出买卖动作
  • 实测显示旧框架下策略波动剧烈,新框架使结果稳定一致
  • 适合研究金融推理、策略设计与因子发现的学者和开发者

大型语言模型(LLMs)在金融领域的应用催生了大量评估基准,从静态知识测试演变为实时交易模拟。然而,现有在线/离线框架普遍忽视一个关键问题:在金融不确定性下的序列决策中,LLMs表现出严重的行为不稳定性。实验表明,当作为交易代理时,即使采用确定性解码,模型仍出现剧烈的运行间差异,动作序列不一致,且相邻时间步常出现非理性的动作突变。这归因于LLMs无状态的自回归特性及其对投资组合分配中连续到离散动作映射的敏感性。这些缺陷从根本上破坏了多数现有基准的可靠性与可复现性。为此,我们提出AlphaForgeBench,将LLMs重新定义为量化研究人员,要求其生成可执行的α因子并构建基于金融知识的因子驱动策略。该范式分离推理与执行机制,实现确定性、可复现的评估,同时契合真实量化研究流程。多组实验验证,该框架有效消除执行引发的不稳定性,为金融推理、策略制定与α发现提供严谨基准。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations. However, existing frameworks for evaluating real-time trading largely overlook a critical failure mode: the severe behavioral instability of LLMs in sequential decision-making under financial uncertainty. Through extensive experiments, we show that when deployed as trading agents, LLMs exhibit extreme run-to-run variance, generate inconsistent action sequences even under deterministic decoding, and frequently produce irrational action flipping across adjacent time steps. We attribute these behaviors to the stateless autoregressive nature of LLMs, which lack persistent memory of prior actions, together with their sensitivity to continuous-to-discrete action mappings in portfolio allocation tasks. These deficiencies fundamentally undermine the reliability and reproducibility of many existing online and offline trading benchmarks. To address these limitations, we propose AlphaForgeBench, a principled evaluation framework that redefines LLMs as quantitative researchers rather than stochastic trading agents. Instead of producing discrete trading actions, AlphaForgeBench requires models to generate executable alpha factors and compose factor-based trading strategies grounded in financial knowledge. This paradigm decouples reasoning from execution mechanics, enabling deterministic and reproducible evaluation while remaining aligned with real-world quantitative research workflows. Extensive experiments across multiple state-of-the-art LLMs demonstrate that AlphaForgeBench eliminates execution-induced instability and provides a rigorous benchmark for evaluating financial reasoning, strategy formulation, and alpha discovery. Webpage at https://finbrain-lab-hkustgz.github.io/AlphaForgeBench

金融AI大模型评测量化研究可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。