用深度强化学习优化加密货币配对交易,提升收益并控制风险。
Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning
- 设计分层筛选与动态均值调整的交易框架,适应高波动市场。
- 在币安合约数据上表现超越传统策略,风险调整后收益显著提升。
- 适合量化交易研究者,尤其关注安全强化学习与加密资产策略者。
本研究探讨深度强化学习(DRL)作为专用执行模块,是否能提升高波动加密货币市场中的配对交易效果。尽管经典配对策略在传统股市表现良好,但在高方差环境中常显僵化且面临严重偏离风险。为此,本文提出新型方法:构建分层‘筛选-排序’配对选择机制,以及专有的‘固定风险、自适应均值’执行模型。系统采用带有LSTM层的近端策略优化(PPO)代理,在严格确定性风险管理边界内决策。基于币安USD-M期货1小时数据的回测显示,优化后的强化学习策略在样本外表现显著优于启发式基准。通过平稳循环块抽样稳健性检验,该策略的风险调整后超额收益在10%水平上统计显著,虽未达5%更严格标准,但仍凸显数字资产极强的特异性波动特征。本研究为量化金融贡献了一种结合统计套利与DRL执行策略的混合架构,并验证了通过统计稳健边界锚定神经策略可有效缓解严重偏离风险,为安全强化学习提供新范式。
原文摘要 · Abstract (English)
This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。