新模型LDA-XGB1兼顾公平性与可解释性,适合金融风控场景
Less Discriminatory Alternative and Interpretable XGBoost Framework for Binary Classification
- 通过双目标优化平衡准确率与公平性,基于分箱和信息值设计
- 在模拟信贷和真实再犯预测数据上均优于传统公平模型
- 保留XGBoost高效性并强制单调约束,结果可解释性强
公平贷款与模型可解释性是金融领域的重要关切,尤其在复杂机器学习模型广泛应用的背景下。为响应消费者金融保护局(CFPB)对防范非法歧视的要求,本文提出LDA-XGB1,一种新型的低歧视性(LDA)机器学习模型,用于公平且可解释的二分类任务。该模型通过双目标优化实现准确率与公平性的平衡,两个目标均基于分箱与信息值构建,并融合XGBoost的预测能力与计算效率,同时保证内在可解释性,包括强制单调性约束。我们在两个数据集上进行评估:SimuCredit(模拟信贷审批数据集)和COMPAS(真实世界再犯预测数据集)。结果表明,LDA-XGB1在预测准确率、公平性与可解释性之间取得有效平衡,通常优于传统公平贷款模型。该方法为金融机构提供了满足公平贷款监管要求的强大工具,同时保留先进机器学习的优势。
原文摘要 · Abstract (English)
Fair lending practices and model interpretability are crucial concerns in the financial industry, especially given the increasing use of complex machine learning models. In response to the Consumer Financial Protection Bureau's (CFPB) requirement to protect consumers against unlawful discrimination, we introduce LDA-XGB1, a novel less discriminatory alternative (LDA) machine learning model for fair and interpretable binary classification. LDA-XGB1 is developed through biobjective optimization that balances accuracy and fairness, with both objectives formulated using binning and information value. It leverages the predictive power and computational efficiency of XGBoost while ensuring inherent model interpretability, including the enforcement of monotonic constraints. We evaluate LDA-XGB1 on two datasets: SimuCredit, a simulated credit approval dataset, and COMPAS, a real-world recidivism prediction dataset. Our results demonstrate that LDA-XGB1 achieves an effective balance between predictive accuracy, fairness, and interpretability, often outperforming traditional fair lending models. This approach equips financial institutions with a powerful tool to meet regulatory requirements for fair lending while maintaining the advantages of advanced machine learning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。