arXiv:2509.22381cs.LG2025-09

用多阶段集成模型提升企业信用风险预测准确率

Enhancing Credit Risk Prediction: A Multi-stage Ensemble Pipeline

  • 融合统计、机器学习与深度学习模型,分阶段优化预测
  • 在2029家美国上市公司数据上实现评级迁移与违约概率精准预测
  • 支持可解释性分析,适合金融风控与决策系统开发

有效的信用风险管理对金融决策至关重要,需构建可靠的模型以预测违约概率并分类金融机构。传统机器学习方法在高维数据、可解释性不足、罕见事件检测及多类风险不平衡方面面临挑战。本研究提出一个全面的多阶段集成框架,整合了有序逻辑回归、有序概率模型等计量经济学方法,以及XGBoost、随机森林、支持向量机、决策树等监督学习算法;包含K近邻等无监督方法、多层感知机等深度学习架构;结合LASSO正则化进行特征选择与降维,并采用纠错输出码作为集成分类器应对不平衡多分类问题。通过每类预测的置换特征重要性分析增强模型透明度。该框架在2,029家美国上市公司的企业信用评级数据集上进行了实证验证,结果表明其显著提升了对企业信用评级变动(升级与降级)及违约概率估计的分类准确性,为战略金融决策支持提供了更精确可靠的计算模型。

原文摘要 · Abstract (English)

Effective credit risk management is fundamental to financial decision-making, requiring robust models to predict default probabilities and classify financial entities. Traditional machine learning approaches face significant challenges when confronted with high-dimensional data, limited interpretability, rare-event detection, and multi-class risk imbalance. This research proposes a comprehensive multi-stage ensemble pipeline that synthesizes multiple complementary models: econometric models including Ordered logit and ordered probit, supervised learning algorithms, including XGBoost, Random Forest, Support Vector Machine, and Decision Tree; unsupervised methods such as K-Nearest Neighbors; deep learning architectures like Multilayer Perceptron; alongside LASSO regularization for feature selection and dimensionality reduction; and Error-Correcting Output Codes as an Ensemble classifier for handling imbalanced multi-class problems. We implement Permutation Feature Importance analysis for each prediction class across all constituent models to enhance model transparency. Our framework can optimize predictive performance while providing a more holistic approach to credit risk assessment. This research contributes to the development of more accurate and reliable computational models for strategic financial decision support by addressing three fundamental challenges in credit risk modeling. The empirical validation of our approach involves analyzing the Corporate Credit Ratings dataset, which contains credit ratings for 2,029 publicly listed US companies. Results demonstrate that our multi-stage ensemble pipeline significantly enhances the accuracy of financial entity classification regarding credit rating migrations (upgrades and downgrades) and default probability estimation.

信用风险集成学习金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。