arXiv:2605.28359cs.AIq-fin.TR2026-05被引 4

新基准让大模型炒股更公平,避免记忆作弊并看清赚钱真因。

From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

论文配图:From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets
图 1 · 摘自论文原文
  • 用数据掩码技术隐藏股票、日期等关键信息,防止模型依赖记忆
  • 分解收益来源,发现多数盈利来自市场趋势而非选股能力
  • 适合评估大模型真实投资能力,尤其关注金融场景中的可迁移技能

评估大语言模型(LLM)在资本市场的盈利能力,常采用端到端交易方式:将模型置于历史市场中进行交易并测量投资组合回报。该方法存在两大缺陷:一是长期回测常与前沿模型的知识截止时间重叠,导致模型通过记忆股票代码、日期、价格和市场叙事来替代真正投资推理;二是原始回报是噪声较大的代理指标,正收益可能源于市场贝塔、风格暴露或有利周期,而非真正的超额收益(alpha)。为此,我们提出KTD-Fin(Knowing-To-Doing Financial Benchmark)——一个端到端的股票市场交易基准,解决上述问题。KTD-Fin采用数据侧掩码协议,一致地隐藏提示和工具中的关键标识符与日历信息,将历史市场记忆与投资决策分离。同时引入类似Barra的绩效归因框架,将投资组合回报分解为市场、风格和选股α三部分。在2024–2026年期间对10个前沿LLM代理在沪深300指数上的测试显示,掩码显著改变了代理的决策逻辑,促使它们转向匿名化的因子驱动推理。归因分析进一步表明,在防泄漏评估下,LLM代理的累计回报主要由被动市场与风格暴露解释,缺乏持续选股α的证据。这说明金融类LLM评估不应仅看是否赚钱,更应关注收益来源是否反映可迁移的投资能力。我们已公开发布KTD-Fin,作为可复现的泄漏控制与归因感知型评估模板。

原文摘要 · Abstract (English)

Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for stock-selection ability, since positive performance may come from market beta, style exposure, or favorable regimes rather than genuine alpha. We introduce KTD-Fin (Knowing-To-Doing Financial Benchmark), an end-to-end stock-market trading benchmark that addresses both issues. KTD-Fin uses a data-side masking protocol to anonymize key identifiers and calendar information consistently across prompts and tools, separating historical market memory from investment decision-making. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. Across ten frontier LLM agents evaluated on the Chinese CSI300 over a 2024--2026 window, masking substantially changes agent rationales, pushing them towards anonymized factor-based reasoning. Attribution analysis further shows that LLM agents' cumulative returns under leakage-controlled evaluation are largely explained by passive market and style exposure, with limited evidence of persistent stock-selection alpha. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also whether the source of returns reflects transferable investment skill. We release KTD-Fin as a reproducible template for leakage-controlled and attribution-aware evaluation of LLM trading agents.

大模型交易金融评估绩效归因防泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。