首个中文岗位技能标注数据集,助力智能招聘匹配。
Chinese-SkillSpan: A Span-Level Dataset for ESCO-Aligned Competency Extraction from Chinese Job Ads

- 基于大模型与人工协同的标注流程,提升效率与精度。
- 涵盖2万+岗位文本,覆盖知识、技能、通用能力等四维度。
- 对标国际标准ESCO,适合招聘算法与职业分析研究者。
职位技能命名实体识别(JobSkillNER)旨在从大规模岗位发布数据中自动提取关键技能信息,对提升人才市场匹配效率和支撑个性化就业服务具有重要意义。据我们所知,本文首次构建了面向中文招聘文本的JobSkillNER数据集。我们制定了适配中文岗位描述的标注规范,并设计了一套大语言模型驱动的宏-微观协同标注流水线:利用大模型进行上下文理解初标,再由专家逐句审核修正。通过该流程,我们在2014-2025年间从四个主流招聘平台收集并标注了超过20,000个实例。基于此,我们发布了Chinese-SkillSpan——首个与ESCO职业能力标准对齐的中文岗位技能数据集,涵盖知识、技能、通识能力与语言能力(LSKT)四个维度。实验表明,该数据集能有效支持模型训练与评估,填补了中文JobSkillNER资源的重大空白,为智能招聘研究提供了重要基准。代码与数据见:https://sites.google.com/view/cn-skillspan-resources。
原文摘要 · Abstract (English)
Job Skill Named Entity Recognition (JobSkillNER) aims to automatically extract key skill information from large-scale job posting data, which is important for improving talent-market matching efficiency and supporting personalized employment services. To the best of our knowledge, this work presents the first Chinese JobSkillNER dataset for recruitment texts. We propose annotation guidelines tailored to Chinese job postings and an LLM-empowered Macro-Micro collaborative annotation pipeline. The pipeline leverages the contextual understanding ability of large language models (LLMs) for initial annotation and then refines the results through expert sentence-level adjudication. Using this pipeline, we annotate more than 20,000 instances collected from four major recruitment platforms over the period 2014-2025. Based on these efforts, we release Chinese-SkillSpan, the first Chinese JobSkillNER dataset aligned with the ESCO occupational skill standard across four dimensions: knowledge, skill, transversal competence, and language competence (LSKT). Experimental results show that the dataset supports effective model training and evaluation, indicating that Chinese-SkillSpan helps fill a major gap in Chinese JobSkillNER resources and provides a useful benchmark for intelligent recruitment research. Code and data are available at https://sites.google.com/view/cn-skillspan-resources .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。