arXiv:2607.10286cs.AIcs.MA2026-07中稿 · EMNLP

测试大模型交易系统能否靠自身智能赚回成本。

Can Agentic Trading Systems Pay for Their Own Intelligence?

论文配图:Can Agentic Trading Systems Pay for Their Own Intelligence?
图 1 · 摘自论文原文
  • 通过追踪交易记录诊断智能决策的收益与成本
  • 发现不同模型有选择资产差、时机判断错等失败模式
  • 适合关注智能交易系统真实盈利性的研究者和投资者

大型语言模型(LLM)代理在交易系统中日益普及,其推理、工具使用和持续决策带来成本,预期产生交易价值。现有评估多关注性能指标,却很少检验代理的可行性:动态决策是否将成本转化为可衡量的增量利润。为此,我们提出TradeLens,一个基于轨迹的诊断工具包,通过交易记录、运行时轨迹和部署配置评估代理交易系统。该工具重建交易路径,将利润与成本归因于可解释证据,并诊断代理是否以及为何为其智能付费。我们在多种骨干模型、资金规模、交易频率和系统架构下进行广泛分析,并讨论部署问题。结果表明,可行性取决于智能到利润的转化:不同模型表现出各异的失败模式,如DeepSeek-V3.2存在资产选择不佳,GLM-4.7出现负向择时;而资金规模、交易频率和架构仅通过放大或削弱决策所附带的择时价值起作用。这一发现将LLM交易代理的评估从能力导向的性能排名转向基于轨迹的智能-利润转化诊断。代码已公开于https://anonymous.4open.science/r/TradeLens。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision-attributed timing value. These findings reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://anonymous.4open.science/r/TradeLens.

智能交易大模型评估利润归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。