arXiv:2509.12456q-fin.TRcs.AI2025-09

用强化学习优化做市策略,应对市场动态变化。

Reinforcement Learning-Based Market Making as a Stochastic Control on Non-Stationary Limit Order Book Dynamics

  • 基于PPO算法构建做市代理,结合真实市场特征建模。
  • 在非平稳市场中表现优于传统方法,收益与风险更优。
  • 适合量化交易研究者和算法做市系统开发者参考。

强化学习为开发自适应、数据驱动的策略提供了有前景的框架,使做市商能够根据与限价订单簿环境的交互优化决策策略。本文探索了将强化学习代理引入做市场景,其中市场动态被显式建模以捕捉真实市场的典型特征,包括聚集的订单到达时间、非平稳的买卖价差与回报漂移、随机的订单量及价格波动性。这些机制旨在提升控制代理的稳定性,并将领域知识融入代理策略学习过程。我们的贡献包括基于近端策略优化(PPO)算法的实际做市代理实现,以及通过基于模拟器的环境对代理在不同市场条件下的表现进行对比评估。分析显示,在财务回报与风险指标上,相比闭式最优解,强化学习代理在非平稳市场条件下仍表现出色,表明所提出的模拟环境可作为训练和预训练强化学习代理的重要工具。

原文摘要 · Abstract (English)

Reinforcement Learning has emerged as a promising framework for developing adaptive and data-driven strategies, enabling market makers to optimize decision-making policies based on interactions with the limit order book environment. This paper explores the integration of a reinforcement learning agent in a market-making context, where the underlying market dynamics have been explicitly modeled to capture observed stylized facts of real markets, including clustered order arrival times, non-stationary spreads and return drifts, stochastic order quantities and price volatility. These mechanisms aim to enhance stability of the resulting control agent, and serve to incorporate domain-specific knowledge into the agent policy learning process. Our contributions include a practical implementation of a market making agent based on the Proximal-Policy Optimization (PPO) algorithm, alongside a comparative evaluation of the agent's performance under varying market conditions via a simulator-based environment. As evidenced by our analysis of the financial return and risk metrics when compared to a closed-form optimal solution, our results suggest that the reinforcement learning agent can effectively be used under non-stationary market conditions, and that the proposed simulator-based environment can serve as a valuable tool for training and pre-training reinforcement learning agents in market-making scenarios.

强化学习做市策略市场建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。