用可解释结构+神经残差,让信贷风险模型既准又透明还公平。
$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

- 将逻辑回归结构与神经网络残差正交分离,兼顾线性可解释性与非线性预测力。
- 在8个公开数据集上,性能接近线性模型,非线性场景下显著优于传统方法。
- 内置诊断工具,帮助判断何时仅看系数足够,何时需关注非线性部分。
信贷风险模型需兼顾预测精度、可解释性与可审计的公平性。逻辑回归虽易解释,但难以捕捉非线性特征;灵活模型预测更好,但解释常为事后附加,无法反映决策规则本身。我们提出 findr(灵活可解释深度回归),一种二元信贷风险建模的半结构化框架,将对数几率分解为可解释的结构成分与正交的神经残差。正交化使系数效应与残差非线性变化分离,训练中引入处理间水桶距离惩罚,通过比较评分分布缓解群体差异。框架还包含诊断工具,用于衡量结构成分对对数几率方差的贡献、决策一致性及局部方向一致性。我们在模拟研究和八个公共信贷数据集上评估 findr,使用评分级准确率-公平性前沿。结果表明:当信号近似线性时,findr 表现接近逻辑回归;在存在非线性结构时,能恢复神经模型的大部分预测优势。诊断工具识别出系数解释与完整模型一致的情况,以及必须考虑残差变化的情形。这些发现支持半结构化建模作为显式权衡性能、公平性与可解释性的实用路径。
原文摘要 · Abstract (English)
Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. Flexible models can improve prediction, but their explanations are often post-hoc and may not describe the decision rule itself. We introduce $\texttt{findr}$, short for flexible, interpretable deep regression, a semi-structured framework for binary credit risk modelling that decomposes the logit into an interpretable structured component and an orthogonal neural residual. The orthogonalisation separates coefficient-based effects from residual nonlinear variation, while an in-processing Wasserstein penalty mitigates group disparities by comparing score distributions during training. The framework also includes diagnostics that measure the structured component's contribution to logit variation, decision agreement, and local directional consistency. We evaluate $\texttt{findr}$ in a simulation study and on eight public credit datasets using score-level accuracy-fairness frontiers. The results show that $\texttt{findr}$ behaves close to logistic regression when the signal is approximately linear, while recovering much of the predictive gain of neural models when nonlinear structure is relevant. The diagnostics identify when coefficient-based explanations remain close to the full fitted model and when residual variation must also be examined. These findings support semi-structured modelling as a practical way to make performance, fairness, and interpretability trade-offs explicit in credit risk decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。