arXiv:2605.16895cs.CEcs.AI2026-05被引 2

警告:大模型交易回报不可信,需过六重验证才可部署

The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence

论文配图:The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
图 1 · 摘自论文原文
  • 提出六项结构有效性测试,检验回报真实性
  • 现有高夏普比率多源于时间污染和参数偏见
  • 建议用大模型做信息接口,独立执行交易决策

端到端大模型交易代理已从研究概念演变为包括FinCon、FinMem、TradingAgents、FinAgent、QuantAgent和FLAG-Trader在内的小型系统生态。部分系统报告的夏普比率极具吸引力,相关基准如FinBen也呈现相似水平。学术界与产业界对架构研究与实际部署之间的差距过度轻视。本文主张:当前报告的超额收益不应被视为可部署证据。在支持部署能力前,必须通过时间完整性、现实摩擦、反事实鲁棒性、预测校准、数值执行和多代理拆解等结构性验证。现有公开证据无法区分稳健预测力与时间污染、未建模摩擦、短窗口夏普不确定性、叙事拟合及参数先验的影响。问题不仅在于评估,更在于结构本身:语言置信度非可交易概率,叙事推理非数值执行,模型先验可能成为未披露的隐含因子暴露。本文贡献最小报告协议套件P1–P6,按主张强度分层适用,并提供保守模块化替代方案——将大模型作为独立校准、风险与执行模块上游的可审计信息接口。代码与复现工具:https://github.com/hj1650782738/Trading。

原文摘要 · Abstract (English)

End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and FLAG-Trader. Several of these report headline Sharpe ratios that would be material if read at face value on a deployment desk, and associated benchmarks such as FinBen report trading-task Sharpe statistics in the same range. The gap between architecture research and deployment claim has been crossed too freely on both sides of the academia--industry divide. We take a position on that gap: reported alpha from end-to-end LLM trading agents should not be treated as deployment evidence. Before such returns can support claims of deployable trading capability, they must survive structural validity tests for temporal integrity, real-world frictions, counterfactual robustness, predictive calibration, numerical execution, and multi-agent disaggregation. Current public evidence cannot yet distinguish robust predictive ability from temporal contamination, unmodeled frictions, short-window Sharpe uncertainty, narrative fitting, and parametric priors. The problem is not only evaluative but structural. Language confidence is not tradable probability, narrative reasoning is not numerical execution, and model priors may become undisclosed implicit factor exposures. We contribute a minimum reporting protocol suite, P1--P6, with tiered applicability by claim strength, and a conservative modular alternative that uses LLMs as auditable information interfaces upstream of independent calibration, risk, and execution modules. Code and reproduction harness: \url{https://github.com/hj1650782738/Trading}.

大模型交易夏普比率可信验证金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。