arXiv:2504.21259cs.CYcs.LG2025-04

STRATA模型通过姓名与地理信息融合,更准确推断种族,减少贷款公平性分析中的偏差。

STRATA: A Name-and-Geography Race Inference Model for Fair Lending and Housing Equity Applications

  • 融合姓名字符与普查区地理信息,用双向LSTM和XGBoost提升推断精度。
  • 将非白人误判为白人的假阳性率从41.8%降至17.8%,在多个数据集上准确率达84.8%以上。
  • 适合用于宏观政策评估,不应用于个体贷款决策,强调使用边界。

在公平贷款合规(ECOA、HMDA、CRA)中,种族与族裔(R&E)的准确推断至关重要,因高达15%的抵押贷款申请缺失种族数据,监管机构需识别差异。现有代理方法如贝叶斯改进姓氏地理编码(BISG)存在与社会经济地位相关的系统性误分类偏差,导致测得差异被低估。本文提出STRATA(Socioeconomic and Tract-Referenced Attribution for Algorithmic analysis),通过堆叠双向LSTM网络整合字符级姓名序列与普查区地理信息,并以XGBoost后处理过滤。核心目标是降低非白人被误判为白人的偏差:STRATA集成模型将该假阳性率从BISG的41.8%降至17.8%。在独立选民登记验证数据集上,基础模型准确率达88.7%,优于单独LSTM(86.4%)、BISG(82.9%)、BIFSG(86.8%)和ZRP(85.8%);集成模型(LSTM+XGBoost)达89.2%。在覆盖全美50州的薪资保护计划贷款验证数据集上,集成模型准确率达84.8%,显著高于仅用姓名的LSTM(76.6%),验证其跨州泛化能力。配套研究将STRATA应用于226万条纽约市房产转让记录。作者警告,此类模型适用于群体层面分析,不应用于个体交易决策。

原文摘要 · Abstract (English)

Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records. Existing proxy methods, including Bayesian Improved Surname Geocoding (BISG), exhibit systematic misclassification biases linked to socioeconomic status that cause measured disparities to understate true levels. This paper introduces STRATA (Socioeconomic and Tract-Referenced Attribution for Algorithmic analysis), a race and ethnicity inference model that integrates character-level name sequences with census tract geolocation via stacked Bidirectional LSTM networks and XGBoost post-filtering. A central goal is reducing the socioeconomically correlated bias that causes non-White individuals to be misclassified as White: STRATA reduces this White False Positive Rate from 41.8% under BISG to 17.8% for the STRATA ensemble. On a held-out voter registration validation dataset, the STRATA base model achieves 88.7% accuracy, outperforming standalone LSTM (86.4%), BISG (82.9%), BIFSG (86.8%), and ZRP (85.8%); the STRATA ensemble (LSTM+XGBoost) reaches 89.2% accuracy. On a national Paycheck Protection Program loan validation dataset covering all 50 states, the STRATA ensemble achieves 84.8% accuracy versus 76.6% for name-only LSTM, confirming cross-state generalizability. A companion paper applies STRATA to 2.26 million New York City residential deed transactions. The authors caution that these models are appropriate for aggregate, population-level analysis and should not be used for individual-level transactional decisions.

种族推断公平贷款地理信息机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。