arXiv:2608.28399cs.AIq-fin.TR2026-08

LLM交易代理的决策存在可预测的负面择时倾向,易被其他参与者利用。

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

  • 通过模拟真实市场环境测试LLM决策模式
  • 多模态、长周期下均发现显著负向择时表现
  • 适合研究市场中可被反向利用的智能体行为

在金融市场中,对价格变动做出系统性反应的序列策略可能被其他市场参与者预测。本文通过RetailAgent实验框架,让大语言模型(LLM)观察匿名化的日内股票价格历史与允许的状态,反复选择持有(long)或空仓(flat),直到下一区间的收益披露。通过消除整体持有比例后,比较同一股票日内路径中持有与空仓区间的收益表现,发现无论模态、时间跨度、状态或模型家族,均存在持续的负向择时。打乱保存的动作序列后该效应显著减弱,表明动作与后续回报间的对齐是负向得分的主要来源。将自生成记忆引入决策会进一步加剧策略持续性,且在同时使用两种动作的股票日中择时更负。这些结果揭示了序列式LLM金融决策中稳定的、可恢复的方向性结构,为研究其他参与者如何应对可预测策略提供了行为信号。

原文摘要 · Abstract (English)

In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.

LLM交易择时偏差行为金融

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。