arXiv:2606.08285cs.AIcs.CE2026-06被引 6

LLM交易研究需统一执行假设,否则结果不可比。

Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

  • 梳理30项研究,检验执行细节的透明度
  • 发现评估假设模糊导致结果难复现
  • 建议制定统一报告标准以提升可比性

大型语言模型(LLMs)和智能体系统在金融交易中的应用日益增多,但各研究间性能难以比较,原因在于数据来源、时间划分、执行时机、换手率处理及交易成本建模等方面的差异。本文对30项与交易相关的主要研究进行了针对性综述与可复现性审计,评估了时点控制、分割透明度、留出评估、成本与换手率处理、执行语义、标的池定义及成果发布等维度。结果显示,尽管模型架构描述相对清晰,但评估所需的执行假设普遍不明确,影响结果的经济可解释性与可复现性。文中通过一个10只股票的案例,仅作为方法论示范,展示显式摩擦与时机选择如何显著压缩主动策略收益。核心结论是:未来LLM交易研究的重点不应仅限于智能体设计,更需建立清晰的执行真实性、可复现性与评估可比性报告标准。

原文摘要 · Abstract (English)

Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling. This article presents a targeted topical review and reproducibility audit of execution realism in LLM-based trading research. A coded evidence matrix covering 30 trade-relevant primary studies is used to assess point-in-time controls, split transparency, held-out evaluation, cost and turnover treatment, execution semantics, universe definition, and artifact release. Across the audited sample, architecture reporting is generally clearer than the evaluation assumptions needed to judge whether a trading result is economically interpretable or reproducible. A 10-equity worked example is included only as a methodological scaffold to illustrate how explicit friction and timing choices can materially compress active-strategy results. The main conclusion is that the next useful step for LLM trading research is not only better agent design, but also clearer reporting standards for execution realism, reproducibility, and evaluation comparability.

LLM交易可复现性评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。