解析动态规划中误差传播机制,提升期权定价精度
Error Propagation in Dynamic Programming: From Stochastic Control to Option Pricing
- 在再生核希尔伯特空间中用非参数回归与蒙特卡洛采样逼近价值函数
- 严格控制每步误差并揭示其从到期日向初始时刻的传播规律
- 为美式期权定价提供理论支持,适合金融工程与强化学习研究者
本文研究离散时间随机最优控制(SOC)的理论与方法基础。在通用动态规划框架下构建控制问题,引入收敛性分析所需数学结构。通过结合非参数回归与蒙特卡洛子采样序列估计关联价值函数:回归步骤在再生核希尔伯特空间(RKHS)中进行,采用经典核岭回归(KRR)算法;蒙特卡洛方法用于估计继续价值。为评估价值函数估计器的准确性,提出自然误差分解,并严格控制每个时间步的误差项。进一步分析误差如何从到期日反向传播至初始阶段,这一方面在现有文献中相对较少关注。最后,说明该分析可自然应用于关键金融场景——美式期权定价。
原文摘要 · Abstract (English)
This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time. We start formulating the control problem in a general dynamic programming framework, introducing the mathematical structure needed for a detailed convergence analysis. The associate value function is estimated through a sequence of approximations combining nonparametric regression methods and Monte Carlo subsampling. The regression step is performed within reproducing kernel Hilbert spaces (RKHSs), exploiting the classical KRR algorithm, while Monte Carlo sampling methods are introduced to estimate the continuation value. To assess the accuracy of our value function estimator, we propose a natural error decomposition and rigorously control the resulting error terms at each time step. We then analyze how this error propagates backward in time-from maturity to the initial stage-a relatively underexplored aspect of the SOC literature. Finally, we illustrate how our analysis naturally applies to a key financial application: the pricing of American options.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。