构建超大规模职业路径数据集,支持从简历文本精准预测职业发展。
KARRIEREWEGE: A Large Scale Career Path Prediction Dataset
- 构建50万+职业路径数据集,链接ESCO职业分类体系。
- 合成岗位标题与描述,提升对非结构化简历的预测性能。
- 适配真实求职场景,助力招聘、职业规划等应用落地。
准确的职业路径预测可为求职者、招聘人员、人力资源及项目管理者提供支持。然而,公开可用的数据与工具仍十分匮乏。本文提出KARRIEREWEGE,一个涵盖超过50万条职业路径的综合性公开数据集,显著超越以往数据规模。我们将其与ESCO职业分类体系关联,为职业轨迹预测提供宝贵资源。针对简历中常见的自由文本输入问题,通过合成岗位标题与描述构建KARRIEREWEGE+,使模型能更准确地从非结构化数据中进行预测,贴近真实应用场景。我们在该数据集及先前基准上评估现有最先进(SOTA)模型,结果显示在自由文本任务下性能与鲁棒性均有提升,归因于合成数据的增强作用。
原文摘要 · Abstract (English)
Accurate career path prediction can support many stakeholders, like job seekers, recruiters, HR, and project managers. However, publicly available data and tools for career path prediction are scarce. In this work, we introduce KARRIEREWEGE, a comprehensive, publicly available dataset containing over 500k career paths, significantly surpassing the size of previously available datasets. We link the dataset to the ESCO taxonomy to offer a valuable resource for predicting career trajectories. To tackle the problem of free-text inputs typically found in resumes, we enhance it by synthesizing job titles and descriptions resulting in KARRIEREWEGE+. This allows for accurate predictions from unstructured data, closely aligning with real-world application challenges. We benchmark existing state-of-the-art (SOTA) models on our dataset and a prior benchmark and observe improved performance and robustness, particularly for free-text use cases, due to the synthesized data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。