arXiv:2507.08019cs.CLecon.GN2025-07被引 2

测试大模型在简历筛选中是否稳定,发现其表现远不如人类专家。

Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks

  • 用三款大模型对比不同公司背景下的简历评估
  • 模型表现受上下文影响大,与人类专家差异显著(p<0.01)
  • GPT最能适应公司背景,但整体仍不靠谱,适合研究者参考

本研究考察大型语言模型(LLMs)在简历筛选任务中是否表现出一致行为(信号)或随机波动(噪声),并评估其性能与人类专家的差距。基于控制数据集,测试了Claude、GPT和Gemini三款模型在四种情境(无公司、企业1[跨国公司]、企业2[初创公司]、简化上下文)下的表现,使用相同及随机排列的简历,并与三位人力资源专家进行对比。方差分析显示,在八种仅限模型的条件下,有四种存在显著均值差异;所有模型与人类评估结果间均呈显著差异(p < 0.01)。配对t检验表明,GPT对公司上下文响应强烈(p < 0.001),Gemini部分适应(企业1时p = 0.038),Claude适应性最弱(p > 0.1)。元认知分析揭示了模型采用的权重策略与人类显著不同。结果表明,尽管大模型在详细提示下可提供可解释的判断模式,但在实际招聘场景中与人类判断存在根本性偏离,为自动化招聘系统部署提供了重要依据。

原文摘要 · Abstract (English)

This study investigates whether large language models (LLMs) exhibit consistent behavior (signal) or random variation (noise) when screening resumes against job descriptions, and how their performance compares to human experts. Using controlled datasets, we tested three LLMs (Claude, GPT, and Gemini) across contexts (No Company, Firm1 [MNC], Firm2 [Startup], Reduced Context) with identical and randomized resumes, benchmarked against three human recruitment experts. Analysis of variance revealed significant mean differences in four of eight LLM-only conditions and consistently significant differences between LLM and human evaluations (p < 0.01). Paired t-tests showed GPT adapts strongly to company context (p < 0.001), Gemini partially (p = 0.038 for Firm1), and Claude minimally (p > 0.1), while all LLMs differed significantly from human experts across contexts. Meta-cognition analysis highlighted adaptive weighting patterns that differ markedly from human evaluation approaches. Findings suggest LLMs offer interpretable patterns with detailed prompts but diverge substantially from human judgment, informing their deployment in automated hiring systems.

大模型评估简历筛选人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。