arXiv:2605.20348q-fin.CPcs.AI2026-05

记忆让AI交易代理在竞合中表现更优,超越理论基准。

Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

论文配图:Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution
图 1 · 摘自论文原文
  • 引入历史反馈与记忆机制,使代理根据实时状态调整策略
  • 有记忆的代理实现更低执行亏损,且优势更持久
  • 适合研究多智能体强化学习在金融交易中的应用

本文研究深度强化学习代理在共享最优执行环境中是否能持续产生超竞争性结果,即实现低于博弈论竞争基准的执行亏损。通过两代理的Almgren-Chriss清仓博弈,分析了代理在获得阶段内环境反馈、价格中间价信息及过往知识时的学习行为。首先使用事前计划代理消除阶段内反馈,以隔离执行前即承诺完整清仓路径的情况;随后引入多种DDQN架构,允许代理基于动态状态进行条件决策。结果表明,当代理可访问阶段内历史信息,特别是近期价格和自身过往动作时,超竞争性结果显著更频繁且更持久。这说明该博弈中的超竞争行为并非源于多智能体学习或仅当前价格观测,而是由反馈、记忆及沿实际执行路径的状态依赖交互共同驱动。

原文摘要 · Abstract (English)

In this paper, we investigate whether deep reinforcement-learning agents interacting in a shared optimal-execution environment can sustain supra-competitive outcomes, in the sense of achieving lower implementation shortfalls than the relevant game-theoretical competitive benchmark. We study a two-agent Almgren-Chriss liquidation game and examine how learned behavior depends on intra-episode environment feedback, the ability to interpret the mid-price and the agent's knoledge of the past. We first use ex-ante schedule-learning agents to remove intra-episode feedback and isolate what can arise when agents commit to complete liquidation trajectories before execution begins. We then allow agents to condition on the evolving state using a variety of DDQN architectures. We find that, when agents are given access to intra-episode history, especially recent prices and own past actions, supra-competitive outcomes become substantially more frequent and more persistent. These findings indicate that supra-competitive behavior in this execution game is driven not by multi-agent learning or by current price observation alone, but by feedback, memory, and state-contingent interaction along the realized execution path.

强化学习交易执行多智能体记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。