arXiv:2507.18680cs.LGcs.AI2025-07被引 1

用强化学习打造能自适应市场的智能做市商

Market Making Strategies with Reinforcement Learning

  • 将做市任务建模为强化学习问题,设计可单/多智能体运行的策略
  • 结合奖励工程与多目标优化,有效平衡库存风险与利润
  • 提出新算法POW-dTS,实现策略动态切换以应对市场变化

本论文系统研究了将强化学习(RL)应用于金融市场的做市问题。做市商虽提供流动性,却面临库存风险、竞争激烈和市场非平稳等挑战。研究将做市任务建模为强化学习问题,设计可在单/多智能体环境下运行的代理,并采用两种互补方法解决库存管理:一是动态奖励塑造,二是多目标强化学习(MORL)的帕累托前沿优化。针对市场非平稳性,提出基于折扣汤普森采样的新策略加权算法POW-dTS,使代理能动态选择并组合预训练策略,持续适应市场变化。实验表明,所提方法在多种性能指标上显著优于传统及基线算法。本研究为构建鲁棒、高效、自适应的做市智能体提供了新方法与洞见,验证了强化学习在复杂金融系统中的变革潜力。

原文摘要 · Abstract (English)

This thesis presents the results of a comprehensive research project focused on applying Reinforcement Learning (RL) to the problem of market making in financial markets. Market makers (MMs) play a fundamental role in providing liquidity, yet face significant challenges arising from inventory risk, competition, and non-stationary market dynamics. This research explores how RL, particularly Deep Reinforcement Learning (DRL), can be employed to develop autonomous, adaptive, and profitable market making strategies. The study begins by formulating the MM task as a reinforcement learning problem, designing agents capable of operating in both single-agent and multi-agent settings within a simulated financial environment. It then addresses the complex issue of inventory management using two complementary approaches: reward engineering and Multi-Objective Reinforcement Learning (MORL). While the former uses dynamic reward shaping to guide behavior, the latter leverages Pareto front optimization to explicitly balance competing objectives. To address the problem of non-stationarity, the research introduces POW-dTS, a novel policy weighting algorithm based on Discounted Thompson Sampling. This method allows agents to dynamically select and combine pretrained policies, enabling continual adaptation to shifting market conditions. The experimental results demonstrate that the proposed RL-based approaches significantly outperform traditional and baseline algorithmic strategies across various performance metrics. Overall, this research thesis contributes new methodologies and insights for the design of robust, efficient, and adaptive market making agents, reinforcing the potential of RL to transform algorithmic trading in complex financial systems.

强化学习做市商算法交易动态策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。