LLM在评估人才时更看重机构名气而非申请人姓名或国籍,且期刊声誉影响更大。
Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals

- 通过三组因子实验,检验了机构声誉、地理来源和姓名种族对LLM评价的影响。
- 机构等级每提升一级,评分高0.297分;期刊声望影响力是机构的5.7倍。
- 发表在《自然》可弥补低名院校背景,尤其对非顶尖院校候选人帮助更大。
我们研究大型语言模型(LLMs)是否在候选人评估中系统性地因申请者姓名的族裔、机构声望及地理来源而产生偏见。报告了三项因子实验(4,320次API调用,四种LLM,五个专业领域)。研究1(3×4设计)发现机构层级存在显著+0.297分(10分制)的梯度效应(95%置信区间:+0.175至+0.422),而姓名起源影响微弱且不显著(95%置信区间跨零)。研究2(2×2 声誉×国家设计)打破声誉与地理的混淆:声誉效应(+0.185;95%置信区间:+0.093至+0.275)比国家起源效应(+0.126;95%置信区间:+0.037至+0.218)高出1.5倍。研究3(2×2 期刊×机构设计)显示,期刊声望(《自然》对比边缘开放获取期刊)的影响是机构声望的5.7倍:期刊效应为+1.937(95%置信区间:+1.811至+2.062),机构效应为+0.341(95%置信区间:+0.184至+0.504)。验证了“挽救效应”:在《自然》发表能更强烈地弥补低声誉机构背景,对瓜亚基尔大学候选人(+2.127)的补偿效果大于麻省理工学院(+1.745)。结果通过中立哲学偏见指数NBI<T,I,F>量化,其中I分量揭示低声誉个体面临更高的评价不一致性,构成一种认知劣势,此点无法仅靠均值指标捕捉。代码与数据:https://github.com/mleyvaz/geo-bias-llm
原文摘要 · Abstract (English)
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI<T,I,F>; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。