信用评分模型误判改进,真实风险却在上升。
The Illusion of Improvement: Reject Inference Strategies in Credit Scoring

- 通过故意放行部分被拒申请人,打破数据反馈循环。
- 仅2%-5%的探索率即可有效识别系统恶化,成本极低。
- 准确率误导判断,拒绝质量才是真实评估标准。
拒绝推理方法广泛用于缓解信用评分中的生存偏差,但其有效性仍不明确。我们系统评估多种方法,发现一种结构性缺陷:在自然重训练周期中,模型准确率提升而召回率下降,造成虚假改进幻觉,导致从业者误认为系统变好,实则拒识能力(正确筛除违约者的能力)正在恶化。为此提出受控探索策略,无需统计假设:贷款方有意识批准部分被拒申请人并观察其真实结果。结果显示,准确率建议不探索,而拒识质量则随探索提升,证实标准评估指标在选择偏差下具有误导性。实验表明,仅2%-5%的探索率即足以以近零成本诊断反馈环严重程度。该结论在两种机器学习方法和三个真实数据集上一致,提示现有评估协议不足以评估存在生存偏差的模型。
原文摘要 · Abstract (English)
Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood. We systematically evaluate several such methods and uncover a structural failure mode: in a natural retraining cycle, models whose accuracy improves while recall collapses create an illusion of improvement that leads practitioners to believe the system is getting better when, in fact, its rejection quality -- the ability to correctly screen out defaulters -- is deteriorating. We then propose a controlled exploration strategy that breaks the feedback loop without statistical assumptions: the lender deliberately approves a fraction of rejected applicants and observes their true outcomes. We show that accuracy and rejection quality give opposite recommendations on whether to explore: accuracy favors no exploration, while rejection quality improves with it, confirming that standard evaluation metrics are misleading under selection bias. Even minimal exploration rates (2--5\%) prove sufficient in our experiments to diagnose the severity of the feedback loop at near-zero cost. Our findings are consistent across two machine learning methods and three real-world datasets, and suggest that standard evaluation protocols are inadequate for assessing models trained under survival bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。