测试大模型对语言细节的偏见,发现含糊表达常被不公平扣分。
I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- 用100组对话模拟招聘场景,控制语义相同只变语言风格。
- 含模糊表达的回答平均得分低25.6%,显示系统存在语言歧视。
- 适合关注AI公平性、招聘自动化和语言偏见的研究者使用。
本文提出一个全面的基准,用于评估大语言模型(LLMs)对语言标记(linguistic shibboleths)的响应,这些微妙的语言特征可能无意中暴露性别、社会阶层或地域背景等人口属性。通过100组经验证的问答对构建的面试模拟,我们发现尽管内容质量相当,但大模型系统性地惩罚某些语言模式,尤其是含糊表达(hedging language)。该基准通过生成在语义上等价但语言特征可控的变体,实现对特定现象的精确测量,从而量化自动化评估系统中的群体偏差。我们在多个语言维度上验证了该方法,结果显示含糊回应的平均评分低25.6%。实验还证明该基准能有效识别模型特有的偏见,为检测和衡量人工智能系统中的语言歧视提供了基础框架,具有广泛应用于公平决策自动化场景的意义。
原文摘要 · Abstract (English)
This paper introduces a comprehensive benchmark for evaluating how Large Language Models (LLMs) respond to linguistic shibboleths: subtle linguistic markers that can inadvertently reveal demographic attributes such as gender, social class, or regional background. Through carefully constructed interview simulations using 100 validated question-response pairs, we demonstrate how LLMs systematically penalize certain linguistic patterns, particularly hedging language, despite equivalent content quality. Our benchmark generates controlled linguistic variations that isolate specific phenomena while maintaining semantic equivalence, which enables the precise measurement of demographic bias in automated evaluation systems. We validate our approach along multiple linguistic dimensions, showing that hedged responses receive 25.6% lower ratings on average, and demonstrate the benchmark's effectiveness in identifying model-specific biases. This work establishes a foundational framework for detecting and measuring linguistic discrimination in AI systems, with broad applications to fairness in automated decision-making contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。