arXiv:2606.14268stat.MLcs.LG2026-06

用梯度提升估计保险尾部风险,提升预测稳定性与准确性

Gradient boosting for extremes: sampling theory and application to insurance

论文配图:Gradient boosting for extremes: sampling theory and application to insurance
图 1 · 摘自论文原文
  • 对广义帕累托分布进行正交重参数化,降低梯度相关性
  • 在18000+医疗过失索赔数据上,发现结案天数是尾部风险主因
  • 给出非渐近误差界,揭示偏差-方差权衡机制

本文为峰值超阈值建模中基于梯度提升估计协变量依赖的广义帕累托(GP)分布,建立了统计学习理论。通过正交重参数化使GP似然的Fisher信息矩阵对角化,将估计问题置于经验风险最小化(ERM)框架下,推导出梯度提升估计器的非渐近误差界。分析考虑了三类误差来源:统计波动、由GP模型渐近性质带来的近似偏差(受二阶正则变化控制),以及有限迭代次数导致的近似误差,明确揭示了偏差-方差权衡。模拟实验表明,该重参数化显著降低训练中的梯度相关性,提高收敛稳定性。方法应用于德克萨斯州保险局提供的医疗过失保险数据集(含超过18,000条已结案索赔),结果表明梯度提升能良好拟合赔付分布尾部,且结案天数是尾部重度的主要预测因子,与早期精算文献结论一致。

原文摘要 · Abstract (English)

We develop a statistical learning theory for gradient boosting applied to the estimation of covariate-dependent Generalized Pareto (GP) distributions in the context of Peaks-over-Threshold modeling. After an orthogonal reparametrization of the GP likelihood that diagonalizes its Fisher information matrix, we cast the estimation problem within the Empirical Risk Minimization (ERM) framework and derive non-asymptotic error bounds for the boosting estimator. Our analysis accounts for three distinct sources of error in the process: statistical fluctuations, the approximation bias inherent to the asymptotic nature of the GP model-controlled under second-order regular variation-and the approximation error associated with the finite number of boosting iterates, making explicit the resulting bias-variance trade-off. We illustrate the practical benefits of the reparametrization through simulations, showing that it significantly reduces gradient correlation during training and improves convergence stability. The methodology is applied to a medical malpractice insurance dataset from the Texas Department of Insurance, comprising over 18 000 closed claims. The gradient boosting approach yields a good fit for the tail of settlement cost distributions and reveals that the number of days to settlement is the dominant predictor of tail heaviness, consistent with earlier findings in the reserving literature.

梯度提升尾部风险保险精算极值统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。