高收入借款人被误判为数据噪声,导致信用违约预测不公平。
Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

- 通过逐步屏蔽特征,识别出收入与利率等三类偏差来源。
- 高收入违约者召回率比低收入者低16.86个百分点。
- 即使隐藏收入信息,贷款金额等代理变量仍导致不公平。
基于大规模消费者借贷数据(LendingClub,N = 1,344,936),研究发现模型置信度筛选机制存在显著人口统计偏差:高收入违约者被错误标记为标签噪声的比例远高于低收入违约者(Cramer's V ≈ 0.03–0.07)。从公平性视角重新审视,高收入与低收入违约者在真正阳性率(召回率)上存在16.86个百分点的差距。采用序列特征屏蔽法揭示三种偏差来源:(1) 直接依赖申请人自报收入;(2) 算法吸收了原始利率中隐含的机构偏见;(3) 即使去除收入和利率后,仍存在3.55个百分点(交叉验证)和2.56个百分点(测试集)的残余差距(Z = -4.04,p < 0.0001)。SHAP分析表明,贷款金额、房产拥有状态等结构性代理变量是维持该差距的关键因素。研究指出,仅屏蔽敏感属性无法保障公平,因制度定价与行为代理变量会重构被隐藏信号。对监管金融机构的数据驱动流程审计具有重要启示。
原文摘要 · Abstract (English)
Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this filtering convention on a large-scale consumer lending sample (LendingClub, N = 1,344,936) uncovers an underlying demographic asymmetry: high-income defaulters are disproportionately classified as label noise relative to low-income defaulters (Cramer's V approximately 0.03-0.07). Re-examining this behavior through the lens of equal opportunity [Hardt et al., 2016] reveals a far more severe discrepancy: a 16.86 percentage point gap in true positive rate (recall) between high- and low-income borrowers who ultimately defaulted. Implementing a sequential feature-blinding methodology allows us to isolate the drivers of this disparity across three distinct mechanisms: (1) direct reliance on self-reported applicant income; (2) algorithmic absorption of upstream institutional bias encoded within origination interest rates; and (3) a residual disparity (3.55 percentage points in cross-validation; 2.56 percentage points on a held-out test partition, Z = -4.04, p < 0.0001) that remains even after purging both income and interest rates from the model. Out-of-sample signed SHAP valuations demonstrate that this residual gap is maintained by structural proxies, most notably loan amount and home ownership status. These empirical findings show that simply blinding an algorithm to sensitive attributes fails to ensure fairness when institutional pricing decisions and behavioral proxy variables collectively reconstruct the omitted signals. We outline the practical implications of these findings for auditing data-centric AI workflows within regulated financial institutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。