LLM投资策略长期表现不及市场,暴露出其在不同行情下的风险控制缺陷。
Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?
- 构建FINSABER框架,在20年跨度、100+股票上验证策略稳健性。
- 实证显示LLM策略在长期内优势消失,牛市保守、熊市激进导致亏损。
- 提示需重视趋势识别与行情感知,而非单纯堆叠模型复杂度。
大语言模型(LLMs)已被用于资产定价和股票交易,使AI能从非结构化金融数据中生成投资决策。然而,多数对基于时序的LLM投资策略评估局限于短期和小规模股票池,易受幸存者偏差和数据窥探偏差影响,高估其有效性。本文提出FINSABER框架,系统性地在更长周期和更大标的范围内评估这类策略。对超过20年、100多个标的的回测结果显示,先前报告的LLM优势在更广泛样本和长期检验下显著减弱。市场状态分析进一步表明,LLM策略在牛市中过于保守,表现低于被动基准;在熊市中又过度激进,造成重大损失。这些发现强调,需开发能优先捕捉趋势并具备行情感知风险控制能力的LLM策略,而非仅依赖框架复杂度的扩展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently been leveraged for asset pricing tasks and stock trading applications, enabling AI agents to generate investment decisions from unstructured financial data. However, most evaluations of LLM timing-based investing strategies are conducted on narrow timeframes and limited stock universes, overstating effectiveness due to survivorship and data-snooping biases. We critically assess their generalizability and robustness by proposing FINSABER, a backtesting framework evaluating timing-based strategies across longer periods and a larger universe of symbols. Systematic backtests over two decades and 100+ symbols reveal that previously reported LLM advantages deteriorate significantly under broader cross-section and over a longer-term evaluation. Our market regime analysis further demonstrates that LLM strategies are overly conservative in bull markets, underperforming passive benchmarks, and overly aggressive in bear markets, incurring heavy losses. These findings highlight the need to develop LLM strategies that are able to prioritise trend detection and regime-aware risk controls over mere scaling of framework complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。