AI做市商学会利用大单交易引发的价格波动获利,加剧中频交易者损失。
When AI Trading Agents Compete: Adverse Selection of Meta-Orders by Reinforcement Learning-Based Market Making
- 用强化学习模拟高频做市商,通过脉冲控制应对订单簿变化。
- 训练后做市商能捕获中频大单引发的价格漂移,实现盈利。
- 适合关注高频交易策略与市场结构的从业者和研究者。
我们研究中频交易者如何被投机性高频交易者不利选择。在霍克斯限价订单簿(Hawkes LOB)模型中,采用强化学习(RL)模拟高频做市商行为。相比传统外生价格冲击模型,霍克斯模型能捕捉内生价格冲击及其他市场特性(Jain et al. 2024a)。鉴于现实中做市商无法对每个订单簿事件实时调整策略,我们采用脉冲控制强化学习框架(Jain et al. 2025)构建高频做市代理。仿真中使用近端策略优化(PPO)与自模仿学习。为复现不利选择现象,测试该RL代理与执行大单的中频交易者(MFT)对抗,结果表明:经训练后,做市代理能有效利用大单引发的价格漂移获利。近期实证研究显示,中频交易者正日益受高频交易者不利选择影响。随着高频交易持续扩张,中频交易者的滑点成本可能进一步上升。然而,我们未观察到做市商利润上升必然导致中频交易者滑点显著增加。
原文摘要 · Abstract (English)
We investigate the mechanisms by which medium-frequency trading agents are adversely selected by opportunistic high-frequency traders. We use reinforcement learning (RL) within a Hawkes Limit Order Book (LOB) model in order to replicate the behaviours of high-frequency market makers. In contrast to the classical models with exogenous price impact assumptions, the Hawkes model accounts for endogenous price impact and other key properties of the market (Jain et al. 2024a). Given the real-world impracticalities of the market maker updating strategies for every event in the LOB, we formulate the high-frequency market making agent via an impulse control reinforcement learning framework (Jain et al. 2025). The RL used in the simulation utilises Proximal Policy Optimisation (PPO) and self-imitation learning. To replicate the adverse selection phenomenon, we test the RL agent trading against a medium frequency trader (MFT) executing a meta-order and demonstrate that, with training against the MFT meta-order execution agent, the RL market making agent learns to capitalise on the price drift induced by the meta-order. Recent empirical studies have shown that medium-frequency traders are increasingly subject to adverse selection by high-frequency trading agents. As high-frequency trading continues to proliferate across financial markets, the slippage costs incurred by medium-frequency traders are likely to increase over time. However, we do not observe that increased profits for the market making RL agent necessarily cause significantly increased slippages for the MFT agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。