arXiv:2506.19789q-fin.RMcs.LG2025-06

对比5种机器学习模型,提升贷款违约概率预测精度

A comparative analysis of machine learning algorithms for predicting probabilities of default

  • 用随机森林、XGBoost等5种算法比对逻辑回归
  • 在信用数据集上,梯度提升类模型表现更优
  • 适合金融风控从业者参考模型选型

预测潜在贷款的违约概率(PD)是金融机构的核心目标。近年来,机器学习(ML)算法在各类预测任务中取得显著成果,但在信用风险分析领域仍应用有限。本文通过对比五种预测模型——随机森林、决策树、XGBoost、梯度提升和AdaBoost——与主流的逻辑回归,在Scheule等人(《信用风险分析:R伴侣》)提供的基准数据集上的表现,揭示了各方法的优劣,为贷款组合中的违约概率预测提供了有效的机器学习算法选择依据。

原文摘要 · Abstract (English)

Predicting the probability of default (PD) of prospective loans is a critical objective for financial institutions. In recent years, machine learning (ML) algorithms have achieved remarkable success across a wide variety of prediction tasks; yet, they remain relatively underutilised in credit risk analysis. This paper highlights the opportunities that ML algorithms offer to this field by comparing the performance of five predictive models-Random Forests, Decision Trees, XGBoost, Gradient Boosting and AdaBoost-to the predominantly used logistic regression, over a benchmark dataset from Scheule et al. (Credit Risk Analytics: The R Companion). Our findings underscore the strengths and weaknesses of each method, providing valuable insights into the most effective ML algorithms for PD prediction in the context of loan portfolios.

信用风险机器学习违约预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。