arXiv:2511.11596cs.LGcs.AI2025-11

解决违约损失预测中代理数据混合导致的模型失效问题

Loss Given Default Prediction Under Measurement-Induced Mixture Distributions: An Information-Theoretic Approach

  • 用信息论方法替代传统分层递归,应对数据混合问题
  • 新方法在1218家企业破产数据上达到r²=0.191,RMSE=0.284
  • 揭示杠杆特征信息量远超规模效应,适用于金融与医疗等场景

违约损失(LGD)建模面临核心数据质量问题:90%的训练数据为破产前资产负债表的代理估计值,而非完整破产程序后的实际回收结果。我们证明,这种混合污染的训练结构会导致递归划分方法系统性失效,随机森林在独立测试集上r²为-0.664(劣于均值预测)。基于香农熵与互信息的信息论方法表现更优,在1,218家企业的破产数据(1980–2023)上实现r²=0.191、RMSE=0.284。分析显示,杠杆相关特征含1.510比特互信息,而规模效应仅贡献0.086比特,与监管机构对规模依赖回收的假设相悖。研究为巴塞尔III下缺乏代表性结果数据的金融机构提供实用指导,并可推广至医疗结局、气候预测与技术可靠性等领域,这些领域因观测周期长而不可避免产生训练数据混合结构。

原文摘要 · Abstract (English)

Loss Given Default (LGD) modeling faces a fundamental data quality constraint: 90% of available training data consists of proxy estimates based on pre-distress balance sheets rather than actual recovery outcomes from completed bankruptcy proceedings. We demonstrate that this mixture-contaminated training structure causes systematic failure of recursive partitioning methods, with Random Forest achieving negative r-squared (-0.664, worse than predicting the mean) on held-out test data. Information-theoretic approaches based on Shannon entropy and mutual information provide superior generalization, achieving r-squared of 0.191 and RMSE of 0.284 on 1,218 corporate bankruptcies (1980-2023). Analysis reveals that leverage-based features contain 1.510 bits of mutual information while size effects contribute only 0.086 bits, contradicting regulatory assumptions about scale-dependent recovery. These results establish practical guidance for financial institutions deploying LGD models under Basel III requirements when representative outcome data is unavailable at sufficient scale. The findings generalize to medical outcomes research, climate forecasting, and technology reliability-domains where extended observation periods create unavoidable mixture structure in training data.

LGD建模信息论金融风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。