用AI生成简历增强匹配效果,显著提升招聘系统精准度。
ConFit v2: Improving Resume-Job Matching using Hypothetical Resume Embedding and Runner-Up Hard-Negative Mining
- 用大模型生成虚拟简历,丰富岗位数据用于对比学习。
- 从无标签数据中挖掘高质量难例,提升模型区分能力。
- 在真实数据集上召回率和排序指标均提升13%以上,适合求职推荐场景。
可靠的简历-职位匹配系统有助于企业从大量简历中推荐合适候选人,帮助求职者找到相关职位。然而,由于求职者仅申请少数职位,简历-职位数据集中的交互标签稀疏。我们提出ConFit v2,改进先前的ConFit方法以应对这一稀疏性问题。通过两种技术增强编码器的对比学习过程:一是利用大语言模型生成虚拟参考简历来扩充岗位数据;二是采用新型难负样本挖掘策略,从无标签简历-职位对中构建高质量难例。我们在两个真实世界数据集上评估ConFit v2,结果表明其优于ConFit及先前方法(包括BM25和OpenAI text-embedding-003),在职位排序和简历排序任务中平均召回率提升13.8%,nDCG提升17.5%。
原文摘要 · Abstract (English)
A reliable resume-job matching system helps a company recommend suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a list of job posts. However, since job seekers apply only to a few jobs, interaction labels in resume-job datasets are sparse. We introduce ConFit v2, an improvement over ConFit to tackle this sparsity problem. We propose two techniques to enhance the encoder's contrastive training process: augmenting job data with hypothetical reference resume generated by a large language model; and creating high-quality hard negatives from unlabeled resume/job pairs using a novel hard-negative mining strategy. We evaluate ConFit v2 on two real-world datasets and demonstrate that it outperforms ConFit and prior methods (including BM25 and OpenAI text-embedding-003), achieving an average absolute improvement of 13.8% in recall and 17.5% in nDCG across job-ranking and resume-ranking tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。