arXiv:2504.05683cs.CLcs.AI2025-04

测试大模型在面试评估中的表现,发现其评分接近真人,但纠错和改进建议能力不足。

Towards Smarter Hiring: Are Zero-Shot and Few-Shot Pre-trained LLMs Ready for HR Spoken Interview Transcript Analysis?

  • 对比10个主流大模型与人类专家,评估面试表现
  • GPT-4 Turbo等模型评分与真人相当,但难识别错误
  • 建议采用人机协同,提升反馈质量

本研究全面分析了包括GPT-4 Turbo、GPT-3.5 Turbo、text-davinci-003、text-babbage-001、text-curie-001、text-ada-001、llama-2-7b-chat、llama-2-13b-chat和llama-2-70b-chat在内的多个预训练大语言模型(LLMs)在模拟人力资源(HR)面试评估中的表现,与专家人类评估者进行对比。研究构建了一个名为HURIT(Human Resource Interview Transcripts)的数据集,包含3,890条来自真实场景的HR面试转录文本。结果表明,尽管这些大模型在评分方面表现出色,与人类专家水平相当,但在识别候选人错误及提供具体可操作的改进建议方面仍存在明显不足。研究指出,当前最先进的预训练大模型尚不适合直接用于自动化的面试评估系统。因此,建议采用人机协同模式,通过人工核查不一致项并优化反馈质量,以实现更可靠的评估效果。

原文摘要 · Abstract (English)

This research paper presents a comprehensive analysis of the performance of prominent pre-trained large language models (LLMs), including GPT-4 Turbo, GPT-3.5 Turbo, text-davinci-003, text-babbage-001, text-curie-001, text-ada-001, llama-2-7b-chat, llama-2-13b-chat, and llama-2-70b-chat, in comparison to expert human evaluators in providing scores, identifying errors, and offering feedback and improvement suggestions to candidates during mock HR (Human Resources) interviews. We introduce a dataset called HURIT (Human Resource Interview Transcripts), which comprises 3,890 HR interview transcripts sourced from real-world HR interview scenarios. Our findings reveal that pre-trained LLMs, particularly GPT-4 Turbo and GPT-3.5 Turbo, exhibit commendable performance and are capable of producing evaluations comparable to those of expert human evaluators. Although these LLMs demonstrate proficiency in providing scores comparable to human experts in terms of human evaluation metrics, they frequently fail to identify errors and offer specific actionable advice for candidate performance improvement in HR interviews. Our research suggests that the current state-of-the-art pre-trained LLMs are not fully conducive for automatic deployment in an HR interview assessment. Instead, our findings advocate for a human-in-the-loop approach, to incorporate manual checks for inconsistencies and provisions for improving feedback quality as a more suitable strategy.

面试分析大模型评估人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。