对比10种梯度提升模型,发现轻量高效且适合高基数分类变量的方案。
From Point to probabilistic gradient boosting for claim frequency and severity prediction
- 统一框架对比点预测与概率预测梯度提升算法
- LightGBM和XGBoostLSS在效率上领先,CatBoost在高基数变量下表现更优
- 首次证明模型准确性和适配性可兼得,适合精算建模场景
梯度提升决策树算法在精算领域应用日益广泛,因其预测性能优于传统广义线性模型。本文以统一符号形式对比了现有全部点预测与概率预测梯度提升方法:GBM、XGBoost、DART、LightGBM、CatBoost、EGBM、PGBM、XGBoostLSS、循环GBM和NGBoost。在五个公开数据集上开展综合数值研究,涵盖不同规模及包含高基数分类变量的索赔频率与损失幅度预测任务。阐述了风险暴露差异在频率模型中的处理方式。基于计算效率、预测性能和模型适配性进行比较:LightGBM和XGBoostLSS在效率上胜出;CatBoost在高基数变量下常提升预测性能;完全可解释的EGBM表现接近黑箱模型。研究发现,模型适配性与预测准确性之间无权衡关系,二者可同时实现。
原文摘要 · Abstract (English)
Gradient boosting for decision tree algorithms are increasingly used in actuarial applications as they show superior predictive performance over traditional generalised linear models. Many enhancements to the first gradient boosting machine algorithm exist. We present in a unified notation, and contrast, all the existing point and probabilistic gradient boosting for decision tree algorithms: GBM, XGBoost, DART, LightGBM, CatBoost, EGBM, PGBM, XGBoostLSS, cyclic GBM, and NGBoost. In this comprehensive numerical study, we compare their performance on five publicly available datasets for claim frequency and severity, of various sizes and comprising different numbers of (high cardinality) categorical variables. We explain how varying exposure-to-risk can be handled with boosting in frequency models. We compare the algorithms on the basis of computational efficiency, predictive performance, and model adequacy. LightGBM and XGBoostLSS win in terms of computational efficiency. CatBoost sometimes improves predictive performance, especially in the presence of high cardinality categorical variables, common in actuarial science. The fully interpretable EGBM achieves competitive predictive performance compared to the black box algorithms considered. We find that there is no trade-off between model adequacy and predictive accuracy: both are achievable simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。