用大模型识别简历中的经验虚报,提升招聘评估准确性
Reading Between the Lines: Classifying Resume Seniority with Large Language Models
- 结合真实与合成简历构建混合数据集,模拟夸大与隐藏资历
- 大模型能有效捕捉隐含的资深度语言线索,识别经验虚报现象
- 为减少自我宣传带来的偏见提供可复用的数据与方法
准确评估简历中候选人的职级是关键但具挑战性的任务,常因经验夸大和模糊自我描述而复杂化。本研究探讨大型语言模型(包括微调的BERT架构)在自动化简历职级分类中的有效性。为严格评估模型表现,我们引入一个混合数据集,包含真实简历与合成生成的高难度样本,用于模拟夸张资质和隐性资深度。利用该数据集,评估大模型检测与职级膨胀相关的细微语言线索及隐含专业能力的能力。研究结果表明,该方法为增强人工智能驱动的候选人评估系统、缓解自我推广语言带来的偏差提供了有前景的方向。数据集已开放供研究社区使用:https://bit.ly/4mcTovt
原文摘要 · Abstract (English)
Accurately assessing candidate seniority from resumes is a critical yet challenging task, complicated by the prevalence of overstated experience and ambiguous self-presentation. In this study, we investigate the effectiveness of large language models (LLMs), including fine-tuned BERT architectures, for automating seniority classification in resumes. To rigorously evaluate model performance, we introduce a hybrid dataset comprising both real-world resumes and synthetically generated hard examples designed to simulate exaggerated qualifications and understated seniority. Using the dataset, we evaluate the performance of Large Language Models in detecting subtle linguistic cues associated with seniority inflation and implicit expertise. Our findings highlight promising directions for enhancing AI-driven candidate evaluation systems and mitigating bias introduced by self-promotional language. The dataset is available for the research community at https://bit.ly/4mcTovt
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。