用实时市场数据测试大模型炒股,发现顶尖模型也会亏钱
Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking
- 构建多智能体系统,接入训练后真实股市数据
- 9个顶级大模型在实时环境均出现净亏损
- 揭示历史回测存在信息泄露,适合实盘研究者参考
大语言模型在金融任务中表现不俗,但在复杂基金投资管理中的实际效果仍缺乏评估。现有基准依赖历史回测,导致模型可通过训练数据中的未来信息实现‘时间旅行’,造成信息泄露和性能虚高。为此,我们提出DeepFund,一个实时基金评测工具,采用多智能体架构,直接接入训练截止日期后的实时股票市场数据,确保评估公平无泄漏。对九个全球领先机构的旗舰大模型在个股分析、投资决策、组合管理与风险控制等维度进行实测,结果表明,即使如DeepSeek-V3和Claude-3.7-Sonnet等先进模型,在实时环境中也产生净交易亏损,凸显当前大模型在主动基金管理中的局限性。代码已开源。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset classification. However, their real-world effectiveness in managing complex fund investment remains inadequately assessed. A fundamental limitation of existing benchmarks for evaluating LLM-driven trading strategies is their reliance on historical back-testing, inadvertently enabling LLMs to "time travel"-leveraging future information embedded in their training corpora, thus resulting in possible information leakage and overly optimistic performance estimates. To address this issue, we introduce DeepFund, a live fund benchmark tool designed to rigorously evaluate LLM in real-time market conditions. Utilizing a multi-agent architecture, DeepFund connects directly with real-time stock market data-specifically data published after each model pretraining cutoff-to ensure fair and leakage-free evaluations. Empirical tests on nine flagship LLMs from leading global institutions across multiple investment dimensions-including ticker-level analysis, investment decision-making, portfolio management, and risk control-reveal significant practical challenges. Notably, even cutting-edge models such as DeepSeek-V3 and Claude-3.7-Sonnet incur net trading losses within DeepFund real-time evaluation environment, underscoring the present limitations of LLMs for active fund management. Our code is available at https://github.com/HKUSTDial/DeepFund.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。