arXiv:2603.06733q-fin.RMcs.AI2026-03被引 1

提出兼顾可靠性与公平性的信贷风险评分框架,应对数据漂移挑战。

Calibrated Credit Intelligence: Shift-Robust and Fair Risk Scoring with Bayesian Uncertainty and Gradient Boosting

  • 融合贝叶斯神经网络与公平约束梯度提升,捕捉不确定性并控制群体差异。
  • 在时序数据漂移下仍保持高精度(AUC-ROC 0.912)与低校准误差(Brier 0.087)。
  • 适合金融风控场景,尤其关注模型稳定性与公平性的实际部署者。

信贷风险评分需支持高风险贷款决策,要求概率估计可靠、群体公平性可保障,且能应对数据分布随时间变化。尽管现代机器学习提升了违约预测精度,但在分布漂移下常出现校准不佳,且未加约束训练易引发不公平结果。本文提出面向部署的校准信用智能(CCI)框架:(i) 使用贝叶斯神经风险评分器捕捉认知不确定性,减少过度自信错误;(ii) 采用公平约束梯度提升模型,在保持强表格性能的同时控制群体差异;(iii) 结合时序感知融合策略与后处理概率校准,稳定后期决策阈值。我们在 Home Credit 信用风险模型稳定性基准上,采用时间一致划分验证,对比 LightGBM、XGBoost、CatBoost、TabNet 及独立贝叶斯神经模型。结果表明,CCI 在区分度、校准性、稳定性与公平性之间达到最佳权衡:AUC-ROC 达 0.912,AUC-PR 为 0.438,召回率@1%假正率(Recall@1%FPR)达 0.509,校准误差低(Brier 分数 0.087,ECE 0.015)。在时间漂移下,其 AUC-PR 下降仅 0.017,群体差异更小(人口均等差距 0.046,机会均等差距 0.037),优于无约束提升模型。结果表明,CCI 能在真实部署条件下生成更准确、可靠且更公平的风险评分。

原文摘要 · Abstract (English)

Credit risk scoring must support high-stakes lending decisions where data distributions change over time, probability estimates must be reliable, and group-level fairness is required. While modern machine learning models improve default prediction accuracy, they often produce poorly calibrated scores under distribution shift and may create unfair outcomes when trained without explicit constraints. This paper proposes Calibrated Credit Intelligence (CCI), a deployment-oriented framework that combines (i) a Bayesian neural risk scorer to capture epistemic uncertainty and reduce overconfident errors, (ii) a fairnessconstrained gradient boosting model to control group disparities while preserving strong tabular performance, and (iii) a shiftaware fusion strategy followed by post-hoc probability calibration to stabilize decision thresholds in later time periods. We evaluate CCI on the Home Credit Credit Risk Model Stability benchmark using a time-consistent split to reflect real-world drift. Compared with strong baselines (LightGBM, XGBoost, CatBoost, TabNet, and a standalone Bayesian neural model), CCI achieves the best overall trade-off between discrimination, calibration, stability, and fairness. In particular, CCI reaches an AUC-ROC of 0.912 and an AUC-PR of 0.438, improves operational performance with Recall@1%FPR = 0.509, and reduces calibration error (Brier score 0.087, ECE 0.015). Under temporal shift, CCI shows a smaller AUC-PR drop from early to late periods (0.017), and it lowers group disparities (demographic parity gap 0.046, equal opportunity gap 0.037) compared to unconstrained boosting. These results indicate that CCI produces risk scores that are accurate, reliable, and more equitable under realistic deployment conditions.

信用评分公平性贝叶斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。