arXiv:2601.11379cs.CLcs.AI2026-01被引 1

评估大模型招聘决策逻辑,发现其权重分配存在群体差异。

Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences

  • 用经济实验法构建合成数据,分析模型对技能经验等要素的权重
  • 模型虽整体公平,但交叉群体中生产力信号权重不同
  • 可复现于人类招聘对比,适合研究算法公平性与人机对齐

通用大语言模型在招聘应用中展现出巨大潜力,因其能处理非结构化文本、权衡多重标准,并从间接绩效信号中推断适配度与能力。然而,模型如何赋予各项属性权重,是否符合经济原理、招聘者偏好或社会规范仍不明确。本文提出一种评估框架,借鉴经济学中分析人类招聘行为的方法,基于欧洲主流在线自由职业平台的真实自由职业者资料与项目描述,构建合成数据集,并采用全因子实验设计,量化大模型在评估自由职业者与项目匹配度时对各项关键因素的权重。研究识别出模型优先考虑的核心生产力信号(如技能与经验),并分析其在不同项目情境及人口子群体间的权重变化。结果表明,尽管模型对少数群体平均无明显歧视,但在交叉身份群体中,生产力信号的实际权重存在差异。最后,本文提出可复现该实验范式于人类招聘者的方案,以评估模型与人类决策的一致性。

原文摘要 · Abstract (English)

General-purpose Large Language Models (LLMs) show significant potential in recruitment applications, where decisions require reasoning over unstructured text, balancing multiple criteria, and inferring fit and competence from indirect productivity signals. Yet, it is still uncertain how LLMs assign importance to each attribute and whether such assignments are in line with economic principles, recruiter preferences or broader societal norms. We propose a framework to evaluate an LLM's decision logic in recruitment, by drawing on established economic methodologies for analyzing human hiring behavior. We build synthetic datasets from real freelancer profiles and project descriptions from a major European online freelance marketplace and apply a full factorial design to estimate how a LLM weighs different match-relevant criteria when evaluating freelancer-project fit. We identify which attributes the LLM prioritizes and analyze how these weights vary across project contexts and demographic subgroups. Finally, we explain how a comparable experimental setup could be implemented with human recruiters to assess alignment between model and human decisions. Our findings reveal that the LLM weighs core productivity signals, such as skills and experience, but interprets certain features beyond their explicit matching value. While showing minimal average discrimination against minority groups, intersectional effects reveal that productivity signals carry different weights between demographic groups.

大模型评估招聘公平性人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。