用编程任务检测大模型在招聘等关键决策中的隐性偏见。
FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks
- 将决策问题转化为编程任务,系统探测模型偏见
- 发现高收入家庭申请人被优先录取的隐蔽偏差
- 适合关注AI公平性与可信决策的研究者
大型语言模型(LLMs)正被广泛应用于招聘、升学等高风险决策,其社会偏见引发严重关切。尽管模型能拒绝显式偏见请求,但偏见可能通过内部规划与推理过程隐性泄露。随着代码成为模型逻辑表达的主要形式,我们提出FairCoder基准,将决策任务转化为编程任务,系统评估模型在就业、教育、医疗领域的偏见,涵盖多种公平性定义。针对模型频繁拒绝请求导致现有指标失效的问题,我们引入FairScore,综合衡量拒绝行为与群体结果多样性。在1000样本数据集上对强大LLM的实验揭示了此前未被充分研究的一致性偏见模式,例如在大学录取中更倾向高收入家庭申请人。研究凸显了将大模型用于决策的风险,并为未来研究提供了全面评估框架。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critical concern. While LLMs are trained to refuse explicitly biased requests, bias can be leaked implicitly during LLM planning and reasoning process. As code becomes the primary medium for LLM internal logic-writing, we introduce FairCoder, a benchmark that frames decision-making as coding tasks to systematically probe LLM bias across employment, education, and healthcare domains, covering multiple fairness definitions. Considering that existing metrics may fail when LLMs frequently refuse the request, we propose FairScore, a metric that jointly captures refusal behavior and group-level outcome diversity. Experiments with a 1k-sample dataset on powerful LLMs reveal consistent and previously underexplored bias patterns, such as prioritizing applicants from high-income families in college admissions. Our findings highlight the risks of deploying LLMs as decision-making agents and provide a comprehensive evaluation framework for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。