arXiv:2505.07653cs.CL2025-05被引 3

构建167万条职业轨迹数据集,助力劳动力市场研究

JobHop: A Large-Scale Dataset of Career Trajectories

  • 用大模型处理简历,提取结构化职业信息
  • 覆盖36万+简历,超167万条工作经历,标准化编码
  • 适合研究职业流动、跳槽趋势及职业路径预测

理解劳动力市场动态对政策制定者、雇主和求职者至关重要,但真实职业轨迹的全面数据稀缺。本文介绍JobHop,一个基于比利时弗拉芒地区公共就业服务机构VDAB提供的匿名简历构建的大规模公开数据集。通过大语言模型(LLMs)处理非结构化简历内容,利用多标签分类模型将其映射至标准化的ESCO职业代码,最终获得超过167万条工作经历,来自36.1万份以上简历,实现了职业信息的结构化与标准化。该数据集可支持劳动力市场流动性分析、职业稳定性研究以及职业中断对职业变迁的影响评估,并推动职业路径预测等数据驱动决策。我们还分析了职位分布、职业中断和职业转换等关键特征,展示了其在劳动力市场研究中的应用价值。

原文摘要 · Abstract (English)

Understanding labor market dynamics is essential for policymakers, employers, and job seekers. However, comprehensive datasets that capture real-world career trajectories are scarce. In this paper, we introduce JobHop, a large-scale public dataset derived from anonymized resumes provided by VDAB, the public employment service in Flanders, Belgium. Utilizing Large Language Models (LLMs), we process unstructured resume data to extract structured career information, which is then normalized to standardized ESCO occupation codes using a multi-label classification model. This results in a rich dataset of over 1.67 million work experiences, extracted from and grouped into more than 361,000 user resumes and mapped to standardized ESCO occupation codes, offering valuable insights into real-world occupational transitions. This dataset enables diverse applications, such as analyzing labor market mobility, job stability, and the effects of career breaks on occupational transitions. It also supports career path prediction and other data-driven decision-making processes. To illustrate its potential, we explore key dataset characteristics, including job distributions, career breaks, and job transitions, demonstrating its value for advancing labor market research.

职业轨迹劳动力市场数据集简历分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。