arXiv:2510.09247q-fin.CPcs.LG2025-10被引 1

用深度强化学习优化标普500期权对冲,比传统方法更抗波动和高成本环境。

Application of Deep Reinforcement Learning to At-the-Money S&P 500 Options Hedging

  • 基于TD3算法训练智能体,从历史数据自动学习对冲策略,无需假设价格模型。
  • 在20年跨市场环境下,其夏普比率和信息比率均优于布莱克-斯科尔斯对冲法。
  • 适合高频交易、高成本或极端行情下的量化对冲场景,尤其关注风险控制者。

本文研究深度强化学习在标普500平价期权对冲中的应用。我们构建了一个基于双延迟深度确定性策略梯度(TD3)算法的智能体,通过历史日内标普500看涨期权价格数据(2004–2024)进行训练,输入包括期权价格、标的资产价格、实值程度、到期时间、已实现波动率和当前对冲头寸等六项变量。采用滚动回测流程训练,实现近17年的样本外评估。将该深度强化学习(DRL)智能体与布莱克-斯科尔斯德尔塔对冲策略在相同时间段对比,使用年化收益、波动率、信息比率和夏普比率等指标评估。进一步测试了不同市场条件下的适应性,并引入交易成本与风险敏感惩罚约束。结果表明,在高波动或高成本环境中,DRL智能体显著优于传统方法,展现出更强鲁棒性与灵活性;但当风险敏感参数升高时性能下降。此外,波动率估计周期越长,结果越稳定。

原文摘要 · Abstract (English)

This paper explores the application of deep Q-learning to hedging at-the-money options on the S\&P~500 index. We develop an agent based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, trained to simulate hedging decisions without making explicit model assumptions on price dynamics. The agent was trained on historical intraday prices of S\&P~500 call options across years 2004--2024, using a single time series of six predictor variables: option price, underlying asset price, moneyness, time to maturity, realized volatility, and current hedge position. A walk-forward procedure was applied for training, which led to nearly 17~years of out-of-sample evaluation. The performance of the deep reinforcement learning (DRL) agent is benchmarked against the Black--Scholes delta-hedging strategy over the same period. We assess both approaches using metrics such as annualized return, volatility, information ratio, and Sharpe ratio. To test the models' adaptability, we performed simulations across varying market conditions and added constraints such as transaction costs and risk-awareness penalties. Our results show that the DRL agent can outperform traditional hedging methods, particularly in volatile or high-cost environments, highlighting its robustness and flexibility in practical trading contexts. While the agent consistently outperforms delta-hedging, its performance deteriorates when the risk-awareness parameter is higher. We also observed that the longer the time interval used for volatility estimation, the more stable the results.

强化学习期权对冲量化交易金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。