用集成学习与SHAP分析,让信用违约预测更准更可解释。
Interpretable Credit Default Prediction with Ensemble Learning and SHAP
- 采用集成学习融合多种算法,提升预测性能。
- 模型在准确率、精确率和召回率上表现优异,尤其应对数据不平衡。
- 通过SHAP揭示外部信用分最重要,增强决策透明度。
本研究聚焦信用违约预测问题,基于机器学习构建建模框架,并在Home Credit数据集上对多种主流分类算法进行对比实验。通过数据预处理、特征工程与模型训练,评估了逻辑回归、随机森林、XGBoost、LightGBM等模型在准确率、精确率和召回率上的表现。结果表明,集成学习方法在预测性能上具有显著优势,尤其在处理特征间复杂非线性关系及数据不平衡问题时表现出强鲁棒性。同时,利用SHAP方法分析特征重要性与依赖关系,发现外部信用评分变量在模型决策中起主导作用,有效提升了模型的可解释性与实际应用价值。研究成果为智能风控系统的建设提供了有力参考与技术支持。
原文摘要 · Abstract (English)
This study focuses on the problem of credit default prediction, builds a modeling framework based on machine learning, and conducts comparative experiments on a variety of mainstream classification algorithms. Through preprocessing, feature engineering, and model training of the Home Credit dataset, the performance of multiple models including logistic regression, random forest, XGBoost, LightGBM, etc. in terms of accuracy, precision, and recall is evaluated. The results show that the ensemble learning method has obvious advantages in predictive performance, especially in dealing with complex nonlinear relationships between features and data imbalance problems. It shows strong robustness. At the same time, the SHAP method is used to analyze the importance and dependency of features, and it is found that the external credit score variable plays a dominant role in model decision making, which helps to improve the model's interpretability and practical application value. The research results provide effective reference and technical support for the intelligent development of credit risk control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。