用迁移学习提升数据少的贷款回收率预测精度。
Transfer Learning for Loan Recovery Prediction under Distribution Shifts with Heterogeneous Feature Spaces

- 设计新型Transformer模型,处理不同特征空间的迁移任务。
- 在真实数据上验证,小样本下预测性能显著优于基线。
- 输出概率分布,适合风险管理者做决策参考。
准确预测回收率(RR)是信用风险管理与监管资本确定的核心。许多贷款组合因违约事件稀少而面临数据稀缺问题。迁移学习(TL)可通过利用相关但数据更丰富的源领域信息缓解此问题,但其效果高度依赖分布偏移程度及源-目标特征空间的异质性。本文提出FT-MDN-Transformer,一种专为异质特征空间下RR预测设计的混合密度表格式Transformer架构,可生成贷款级点估计与组合级预测分布,支持多样应用场景。我们在受控蒙特卡洛模拟中系统测试了协变量、条件与标签偏移的影响,并在真实场景中以全球信贷数据(GCD)为源,新债券数据为靶,进行迁移实验。结果表明,当目标数据有限时,该模型优于基准方法,尤其在协变量与条件偏移下表现突出;标签偏移仍具挑战性。此外,其概率预测能紧密匹配实际回收分布,提供比传统点预测更丰富的信息。研究显示,具备分布感知能力的迁移架构可有效提升数据稀缺信贷组合中的回收率预测能力,为异构数据环境下风险管理者提供实用洞见。
原文摘要 · Abstract (English)
Accurate forecasting of recovery rates (RR) is central to credit risk management and regulatory capital determination. In many loan portfolios, however, RR modeling is constrained by data scarcity arising from infrequent default events. Transfer learning (TL) offers a promising avenue to mitigate this challenge by exploiting information from related but richer source domains, yet its effectiveness critically depends on the presence and strength of distributional shifts, and on potential heterogeneity between source and target feature spaces. This paper introduces FT-MDN-Transformer, a mixture-density tabular Transformer architecture specifically designed for TL in RR forecasting across heterogeneous feature sets. The model produces both loan-level point estimates and portfolio-level predictive distributions, thereby supporting a wide range of practical RR forecasting applications. We evaluate the proposed approach in a controlled Monte Carlo simulation that facilitates systematic variation of covariate, conditional, and label shifts, as well as in a real-world transfer setting using the Global Credit Data (GCD) loan dataset as source and a novel bonds dataset as target. Our results show that FT-MDN-Transformer outperforms baseline models when target-domain data are limited, with particularly pronounced gains under covariate and conditional shifts, while label shift remains challenging. We also observe its probabilistic forecasts to closely track empirical recovery distributions, providing richer information than conventional point-prediction metrics alone. Overall, the findings highlight the potential of distribution-aware TL architectures to improve RR forecasting in data-scarce credit portfolios and offer practical insights for risk managers operating under heterogeneous data environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。