arXiv:2510.02209cs.LGcs.CL2025-10被引 32

测试大模型能否在真实股市中赚钱,发现多数表现不如简单持有。

StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?

  • 用真实市场数据训练大模型做买卖决策,模拟多月交易流程。
  • 多数模型收益低于买入持有策略,但少数表现更优且风险控制更好。
  • 适合研究金融智能体、大模型实战能力的学者和从业者。

大型语言模型(LLMs)展现出作为自主代理的强大潜力,具备推理、工具使用和序列决策能力。然而,尽管金融领域具有重要经济价值和复杂的推理需求,现有基准大多局限于静态问答任务,未能反映真实市场的动态特性。为此,我们提出STOCKBENCH——一个无污染的基准,用于评估LLM代理在真实、多月股票交易环境中的表现。代理每日接收价格、基本面和新闻等市场信号,并做出买入、卖出或持有决策。性能通过累积收益、最大回撤和Sortino比率等金融指标衡量,兼顾收益与风险管理。我们评估了多种前沿专有及开源大模型,结果表明多数模型无法超越简单的买入并持有基线,但部分模型展现出更高的收益与更强的风险管理能力。这揭示了基于大模型的交易代理面临的挑战与机遇,说明静态金融问答表现优异并不意味着能有效交易。我们开源STOCKBENCH,以推动未来大模型驱动的金融代理研究。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong potential as autonomous agents, with promising capabilities in reasoning, tool use, and sequential decision-making. While prior benchmarks have evaluated LLM agents in various domains, the financial domain remains underexplored, despite its significant economic value and complex reasoning requirements. Most existing financial benchmarks focus on static question-answering, failing to capture the dynamics of real-market trading. To address this gap, we introduce STOCKBENCH, a contamination-free benchmark designed to evaluate LLM agents in realistic, multi-month stock trading environments. Agents receive daily market signals -- including prices, fundamentals, and news -- and make sequential buy, sell, or hold decisions. Performance is measured using financial metrics such as cumulative return, maximum drawdown, and the Sortino ratio, capturing both profitability and risk management. We evaluate a wide range of state-of-the-art proprietary and open-source LLMs. Surprisingly, most models struggle to outperform the simple buy-and-hold baseline, while some models demonstrate the potential to achieve higher returns and stronger risk management. These findings highlight both the challenges and opportunities of LLM-based trading agents, showing that strong performance on static financial question-answering do not necessarily translate into effective trading behavior. We release STOCKBENCH as an open-source benchmark to enable future research on LLM-driven financial agents.

大模型量化交易金融智能体实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。