arXiv:2605.25527cs.LGcs.CE2026-05

用订单流信号与分组优化策略提升高频交易收益

DeepSeekMath Meets Order Book: Group-Aware Policy Optimization for High-Frequency Directional Trading

论文配图:DeepSeekMath Meets Order Book: Group-Aware Policy Optimization for High-Frequency Directional Trading
图 1 · 摘自论文原文
  • 基于订单流构建状态,用改进的PPO算法优化交易策略
  • 在亚马逊等三只股票上,平均盈亏提升12.3%,回撤降低18.6%
  • 适合关注高频量化交易与强化学习融合的研究者

本文研究将强化学习应用于限价订单簿高频交易,采用基于订单流的状态建模与策略梯度方法。不同于传统的基于价值的RL方法(如表格Q-learning),该方法使用基于策略的方法,如原始PPO及受DeepSeekMath启发的GRPO和GSPO,它们采用分组归一化更新和下行风险感知奖励设计。在简化回测框架下,以价差缩放奖励为基准,对AMZN、AAPL和GOOG三只金融资产进行测试,新策略在净平均盈利(PnL)、盈利能力和最大回撤方面均优于Q-learning基线。结果表明:(1) 订单流信号作为策略强化学习的状态是充分有效的;(2) 分组感知的PPO代理比基于价值的基线更优。

原文摘要 · Abstract (English)

This paper studies reinforcement learning for high-frequency trading on limit order books by pairing an Order-Flow-based state model with policy-gradient methods. Instead of value-based RL techniques like tabular Q-learning, our approach deploys policy-based methods like vanilla PPO and DeepSeekMath-inspired variants like GRPO and GSPO, that use group-normalized updates and downside-aware shaping. On backtests with financial assets AMZN, AAPL, and GOOG under a simplified backtesting setup based on spread-scaled rewards, these new policies improve net average PnL, profitability, and drawdown over the Q-Learning baseline. Our results show that (1) Order-Flow signals are an adequate state for policy RL and (2) group-aware PPO surrogates are preferable over value-based baselines.

高频交易强化学习订单簿策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。