arXiv:2501.07663cs.CL2025-01被引 2

用大模型从120万职位信息中精准提取隐含用工特征。

Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning

  • 结合语义分块与RAG,微调DistilBERT识别复杂岗位特征
  • 对非薪资薪酬、远程工作等关键变量识别准确率显著提升
  • 适合人力资源、招聘平台做数据洞察与智能匹配

本文探索利用大型语言模型(LLMs)从非结构化职位发布信息中提取精细且复杂的岗位特征。基于AdeptID提供的120万条职位数据,我们构建了稳健的处理流程,用于识别和分类远程工作可用性、薪酬结构、教育要求及工作经验偏好等变量。方法结合语义分块、检索增强生成(RAG)与DistilBERT模型微调,克服了传统解析工具在处理复杂或隐含信息时的局限性。通过该技术,显著提升了对非薪资薪酬、推断型远程工作类别等常被误标或遗漏变量的识别能力。我们对微调模型进行了全面评估,分析其优势、局限及可扩展性。本研究展示了LLMs在劳动力市场分析中的潜力,为更精确、可操作的职位数据分析提供了基础。

原文摘要 · Abstract (English)

This paper explores the application of large language models (LLMs) to extract nuanced and complex job features from unstructured job postings. Using a dataset of 1.2 million job postings provided by AdeptID, we developed a robust pipeline to identify and classify variables such as remote work availability, remuneration structures, educational requirements, and work experience preferences. Our methodology combines semantic chunking, retrieval-augmented generation (RAG), and fine-tuning DistilBERT models to overcome the limitations of traditional parsing tools. By leveraging these techniques, we achieved significant improvements in identifying variables often mislabeled or overlooked, such as non-salary-based compensation and inferred remote work categories. We present a comprehensive evaluation of our fine-tuned models and analyze their strengths, limitations, and potential for scaling. This work highlights the promise of LLMs in labor market analytics, providing a foundation for more accurate and actionable insights into job data.

大模型应用岗位分析特征提取人力资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。