arXiv:2601.01709q-fin.PRcs.LG2026-01被引 3

用强化学习优化期权对冲,兼顾风险与成本,提升实际对冲效果。

Reinforcement Learning for Option Hedging: Static Implied-Volatility Fit versus Shortfall-Aware Performance

  • 在黑斯科尔斯框架中引入风险厌恶和交易成本,改进Q学习算法。
  • 动态对冲表现更优,降低亏损概率,优于传统静态定价方法。
  • 适合量化交易、金融工程研究者,关注真实市场表现的模型设计。

我们扩展了黑斯科尔斯框架下的Q-learner(QLBS),引入风险厌恶和交易成本,提出一种新的期权定价复制学习(RLOP)方法。两种方法均兼容标准强化学习算法,且能处理市场摩擦。基于SPY和XOP期权数据,在静态和动态维度评估性能。自适应QLBS在隐含波动率空间中实现更高静态定价精度,而RLOP通过降低亏损概率,在动态对冲中表现更优。结果表明,评估期权定价模型不应仅关注静态拟合,更应重视实际对冲表现。

原文摘要 · Abstract (English)

We extend the Q-learner in Black-Scholes (QLBS) framework by incorporating risk aversion and trading costs, and propose a novel Replication Learning of Option Pricing (RLOP) approach. Both methods are fully compatible with standard reinforcement learning algorithms and operate under market frictions. Using SPY and XOP option data, we evaluate performance along static and dynamic dimensions. Adaptive-QLBS achieves higher static pricing accuracy in implied volatility space, while RLOP delivers superior dynamic hedging performance by reducing shortfall probability. These results highlight the importance of evaluating option pricing models beyond static fit, emphasizing realized hedging outcomes.

强化学习期权对冲量化金融风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。