构建大学课程与职业能力匹配数据集,助力教育与就业衔接
UniSkill: A Dataset for Matching University Curricula to Professional Competencies
- 基于欧洲职业能力分类体系,构建课程与技能匹配数据集
- 用BERT模型实现课程到技能匹配,F1达87%
- 适合教育规划、职业推荐系统研究者使用
技能提取与推荐系统已从招聘方、求职者和教育机构视角被广泛研究。尽管人工智能在职位广告中的应用备受关注,但教学端技能数据的缺乏仍是挑战。本文发布了一个公开数据集,包含人工标注和合成的大学课程与欧洲技能、能力、资格与职业分类(ESCO)中系统分析师及管理组织分析师职业类别下的技能匹配数据,并提供标注指南。具体在两个粒度上进行匹配:课程标题与技能、课程描述句子与技能。我们在该数据集上训练语言模型,为课程-技能与技能-课程匹配的检索与推荐系统提供基线。在部分标注数据上评估模型,所用BERT模型达到87%的F1分数,证明课程与技能匹配是可行任务。
原文摘要 · Abstract (English)
Skill extraction and recommendation systems have been studied from recruiter, applicant, and education perspectives. While AI applications in job advertisements have received broad attention, deficiencies in the instructed skills side remain a challenge. In this work, we address the scarcity of publicly available datasets by releasing both manually annotated and synthetic datasets of skills from the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy and university course pairs and publishing corresponding annotation guidelines. Specifically, we match graduate-level university courses with skills from the Systems Analysts and Management and Organization Analyst ESCO occupation groups at two granularities: course title with a skill, and course sentence with a skill. We train language models on this dataset to serve as a baseline for retrieval and recommendation systems for course-to-skill and skill-to-course matching. We evaluate the models on a portion of the annotated data. Our BERT model achieves 87% F1-score, showing that course and skill matching is a feasible task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。