arXiv:2608.10089stat.APcs.AI2026-08

模型对姓氏的隐含偏见不等于实际决策偏差,两者可能毫无关联。

Status Association Does Not Reliably Predict Decision Leakage

论文配图:Status Association Does Not Reliably Predict Decision Leakage
图 1 · 摘自论文原文
  • 用智利姓氏做社会经济探针,分离潜在关联与实际决策。
  • 多数模型中姓氏偏见未转化为真实选拔中的优势,效果接近零。
  • 提醒评估应直接测量偏见如何影响行为,而非仅看隐含关联。

偏见评估常从模型编码社会关联的证据,直接推断其会影响关键决策。我们以智利姓氏为受控的社会经济探针,测试这一推论是否成立。在8个冻结模型-提供者组合上,每个模型处理1,032个提示,共获得8,256条经验证的原始响应。实验设计将强制的潜在关联与匹配的关键决策(学术选拔、职业招聘、研究资助、法律援助申请)分离开来。在八个模型中有七个显示精英编码姓氏的高地位概率质量高于普通姓氏,所有八种模型中均高于低频对照组。然而,多数系统中精英姓氏与普通姓氏的决策差异接近零。五个模型在预设(±0.10)标准差范围内无显著差异,其余三个结果模糊或边缘,且无一致的精英优势。潜在关联强度无法可靠预测决策泄露(r = 0.201, p = 0.633),跨姓名对-模型单元亦然(r = 0.065, p = 0.565)。核心发现是:潜在社会关联与实际后果性对待在实证上是可分离的构念。评估应直接测量从关联到行动的转化过程。

原文摘要 · Abstract (English)

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.

偏见评估决策公平性模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。