EBM模型兼顾保险赔付预测精度与可解释性,适合风控决策。
Explainable Boosting Machine for Predicting Claim Severity and Frequency in Car Insurance
- 结合GAM与梯度提升,构建可解释的透明预测模型。
- 在真实数据上表现优于主流保险定价模型,且校准更优。
- 能精确分解预测结果至各因素贡献,助力业务洞察。
随着机器学习和深度学习技术的快速发展,精算师和保险行业始终面临预测准确性与可解释性之间的权衡。本文在车险场景下对可解释梯度提升机(EBM)进行了全面的应用评估,聚焦于索赔频率与严重程度建模。EBM将广义加性模型(GAM)的加性结构与循环梯度提升算法相结合,形成一个从设计上就可解释的‘玻璃箱’模型。基于真实世界数据,我们实证展示了其实际应用价值,并与非寿险定价中常用的现代基准模型进行对比。评估涵盖:(i) 外样本预测精度,包括Murphy图和Bregman主导性检验;(ii) 校准性能,采用T-可靠性图和Murphy得分分解。最后,我们揭示了EBM预测与Shapley值之间的关联,证明预测结果可被透明地分解为精确的主效应和成对交互效应,提供超越预测性能的可行动洞察。
原文摘要 · Abstract (English)
With the rapid development of machine learning and deep learning techniques, actuaries and the broader insurance industry face a persistent trade-off between predictive accuracy and interpretability. This paper provides a comprehensive applied assessment of Explainable Boosting Machines (EBM) in a car insurance framework, focusing on claim frequency and severity modeling. EBM combines the additive structure of generalized additive models (GAM) with a cyclic gradient boosting algorithm, resulting in a glass-box model whose predictions are interpretable by design. Using real-world data, we empirically illustrate its practical relevance and compare EBM with modern benchmark models used in non-life insurance pricing. The evaluation considers (i) out-of-sample predictive accuracy, including Murphy diagrams and Bregman dominance tests, and (ii) calibration assessment using T-reliability diagrams and Murphy's score decomposition. Finally, we highlight the link between EBM predictions and Shapley values, showing how predictions can be transparently decomposed into exact main and pairwise interaction effects, providing actionable insights beyond predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。