arXiv:2605.18147cs.LG2026-05被引 2

用预训练模型提升小数据下的信贷风险预测,效果显著。

Foundation Models for Credit Risk Prediction: A Game Changer?

论文配图:Foundation Models for Credit Risk Prediction: A Game Changer?
图 1 · 摘自论文原文
  • 用跨领域数据预训练的表格模型,无需调参直接使用。
  • 在小样本场景下,违约概率与损失率预测准确率明显提升。
  • 适合中小企业贷款等数据少、失衡严重的风控场景。

预测模型在信用风险管理中至关重要,直接影响违约概率和损失的估算。尽管梯度提升模型搭配SHAP解释器已成为行业准标准,但风险模型的持续优化仍是核心目标。与此同时,大语言模型等AI技术快速发展,基础模型凭借大规模跨领域预训练展现出强大潜力。尽管在自然语言和视觉领域已广泛应用,表格数据的基础模型才刚出现。我们推测,在小数据场景(如中小企业贷款或特定企业组合)中,跨领域预训练有助于缓解低违约率和类别不平衡等长期难题。本文对近期提出的表格基础模型进行了全面评估,对比了多种主流机器学习方法,在违约概率(PD)和违约损失率(LGD)两个任务上覆盖多个数据集、指标和实验条件。结果表明,表格基础模型在各类数据集和任务中表现最佳,尤其在数据量减少时提升尤为显著。该成果令人瞩目,因模型为开箱即用,无需超参数调优,兼具易用性与低计算成本。

原文摘要 · Abstract (English)

Predictive models play a pivotal role in credit risk management, guiding critical decisions through accurate estimation of default probabilities and losses. Extensive research has introduced new modeling techniques, complemented by large-scale benchmarking studies consolidating the state-of-the-art. Today, quasi-standards such as gradient-boosting models paired with SHAP explainers have emerged, yet continuous improvement of risk models remains a top priority. Concurrently, rapid advancements in AI, most notably large language models, have disrupted predictive modeling paradigms. Foundation models, pretrained on extensive datasets from diverse domains, have demonstrated remarkable performance by leveraging prior knowledge. While prevalent in natural language processing and computer vision, foundation models for tabular data have only recently emerged. We conjecture that pretraining on out-of-domain data is particularly beneficial in small-data settings, such as SME lending or specialized corporate portfolios, and may help address longstanding challenges including low default portfolios and class imbalance. This paper benchmarks recently proposed tabular foundation models against a broad set of competitors, including established and advanced machine learning techniques, across two core tasks: PD and LGD modeling. Our evaluation encompasses various datasets, performance indicators, and experimental conditions. We find that tabular foundation models generally perform best across datasets and tasks. Moreover, they offer significant improvement in predictive performance as dataset size shrinks. These results are remarkable given that the models are tested out-of-the-box, without hyperparameter tuning, ensuring ease of use and mitigating computational costs.

信贷风险基础模型小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。