用强化学习优化大单交易,自动适应订单簿变化
Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution
- 不依赖预设模型,用数据驱动方式学习交易策略
- 在多种配置下表现优于传统方法,降低市场冲击和执行成本
- 适合量化交易、算法交易研究者参考
我们研究了强化学习在大单交易最优执行中的应用,目标是在长时间内以最小化实现缺口和市场影响的方式逐步完成大额订单。不同于传统的参数化价格动态与影响建模,本文采用无模型、数据驱动的框架。由于策略优化需要反事实反馈,历史数据无法提供,因此使用队列响应模型生成包含瞬时价格影响、非线性及动态订单流响应的真实且可计算的限价单簿模拟。方法上,我们在包含时间、持仓、价格和深度变量的状态空间中训练双深度Q网络(Double Deep Q-Network)智能体,并与现有基准进行对比。数值仿真结果表明,该智能体学习到的策略兼具战略性和战术性,能有效适应订单簿状态,在多个训练配置下均优于标准方法。这些发现有力证明,无模型强化学习可为最优执行问题提供自适应且鲁棒的解决方案。
原文摘要 · Abstract (English)
We investigate the use of Reinforcement Learning for the optimal execution of meta-orders, where the objective is to execute incrementally large orders while minimizing implementation shortfall and market impact over an extended period of time. Departing from traditional parametric approaches to price dynamics and impact modeling, we adopt a model-free, data-driven framework. Since policy optimization requires counterfactual feedback that historical data cannot provide, we employ the Queue-Reactive Model to generate realistic and tractable limit order book simulations that encompass transient price impact, and nonlinear and dynamic order flow responses. Methodologically, we train a Double Deep Q-Network agent on a state space comprising time, inventory, price, and depth variables, and evaluate its performance against established benchmarks. Numerical simulation results show that the agent learns a policy that is both strategic and tactical, adapting effectively to order book conditions and outperforming standard approaches across multiple training configurations. These findings provide strong evidence that model-free Reinforcement Learning can yield adaptive and robust solutions to the optimal execution problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。