arXiv:2607.11715cs.CL2026-07

从44万份简历中自动提取职业轨迹,构建大规模真实文本数据集。

JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes

  • 用可控推理的LLM pipeline从非结构化简历中提取职业路径
  • 产出35.5万条带职业编码、时间与教育等级标注的轨迹数据
  • 适合做劳动力市场分析、求职推荐系统的研究者使用

大规模、丰富标注的职业轨迹数据是人力资源规划、职位推荐和劳动力市场分析的基础,但现有公开数据集或规模小、或无法独立使用,或基于经过标准化的职业代码且由大模型合成而非真实自由文本。我们提出JobHop v2,是公开可用的JobHop数据集的升级版本,通过端到端大语言模型(LLM)从约44万份由弗拉芒公共就业服务局(VDAB)提供的匿名化多语言简历中提取信息。释放的数据集包含355,315条职业轨迹,标注了ESCO职业代码、季度级时间信息以及五级制教育水平,显著提升了原始版本的覆盖范围与标注丰富度。相比v1,JobHop v2引入了基于推理控制的LLM抽取流程与重试机制(实现100% JSON解析率)、更丰富的抽取模式和新的评估协议,该协议以三个互补的人工标注基准为参照。在这些基准上评估,最佳抽取器的表现最接近人工标注者间的一致性上限,差距仅1.1-2.7个百分点。数据集与代码已公开,支持可复现的职业轨迹研究。

原文摘要 · Abstract (English)

Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We present JobHop~v2, an improved version of the publicly available JobHop dataset, constructed through end-to-end large language model (LLM) extraction from a corpus of ${\sim}440{,}000$ pseudonymized, multilingual resumes provided by VDAB, the Flemish Public Employment Service. The released dataset comprises $355{,}315$ career trajectories annotated with ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment, broadening both the coverage and the annotation richness of the original release. Relative to v1, JobHop~v2 introduces a redesigned extraction pipeline based on reasoning-controlled LLM inference with a retry mechanism (achieving a 100% JSON parse rate), a richer extraction schema, and a revised evaluation protocol scored against three complementary annotation baselines. Evaluated against these baselines, our best extractor comes closest to the inter-annotator agreement ceiling among all compared models, trailing it by only 1.1-2.7 percentage points. The dataset and code are publicly released to support reproducible career-trajectory research.

职业轨迹大模型应用劳动力市场数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。