arXiv:2510.19167cs.CL2025-10

测试大模型能否通过科技公司真实招聘评估,结果全军覆没。

"You Are Rejected!": An Empirical Study of Large Language Models Taking Hiring Evaluations

  • 用顶尖大模型模拟求职者回答专业测评题
  • 所有模型均未达到公司标准答案要求
  • 揭示大模型在真实职场评估中的能力短板

随着互联网普及和人工智能快速发展,头部科技公司每年需从数千名申请者中高效筛选高潜力软件与算法工程师。为此,企业建立了多阶段选拔流程,其中包含评估岗位特定能力的标准化招聘测评。受大语言模型在编码与推理任务中表现出色的启发,本文探究一个重要问题:大语言模型能否成功通过此类招聘评估?为此,我们对一份广泛使用的专业测评问卷进行了全面考察,采用前沿大模型生成答案并进行评估。结果出人意料:所有测试的大模型均未能通过招聘评估,其生成答案与公司参考答案存在显著不一致。研究发现表明,尽管大模型具备强大能力,但在真实职场评估中仍无法达到合格标准。

原文摘要 · Abstract (English)

With the proliferation of the internet and the rapid advancement of Artificial Intelligence, leading technology companies face an urgent annual demand for a considerable number of software and algorithm engineers. To efficiently and effectively identify high-potential candidates from thousands of applicants, these firms have established a multi-stage selection process, which crucially includes a standardized hiring evaluation designed to assess job-specific competencies. Motivated by the demonstrated prowess of Large Language Models (LLMs) in coding and reasoning tasks, this paper investigates a critical question: Can LLMs successfully pass these hiring evaluations? To this end, we conduct a comprehensive examination of a widely used professional assessment questionnaire. We employ state-of-the-art LLMs to generate responses and subsequently evaluate their performance. Contrary to any prior expectation of LLMs being ideal engineers, our analysis reveals a significant inconsistency between the model-generated answers and the company-referenced solutions. Our empirical findings lead to a striking conclusion: All evaluated LLMs fails to pass the hiring evaluation.

大模型评估招聘测评真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。