arXiv:2503.03932cs.CL2025-03

构建首个西班牙语技能分类数据集,助力招聘智能化

Tec-Habilidad: Skill Classification for Bridging Education and Employment

  • 创建西班牙语简历技能提取与分类数据集
  • 标注区分知识、技能、能力三类要素
  • 提供深度学习基线模型,适配教育就业衔接场景

近年来,随着技术进步和企业运营方式变革,求职申请与评估流程已发生显著演变。技能提取与分类仍是现代招聘中的关键环节,能更客观地评估候选人并自动匹配岗位需求。但有效评估技能需识别简历中多样化的表达形式,包括直接提及、隐含表述、同义词、缩写、短语及熟练度,并区分硬技能与软技能。尽管大模型(LLMs)在提取和分类技能方面有所助益,但缺乏针对西班牙语求职材料的综合性数据集来评估模型性能。这一空白限制了对模型可靠性和准确性的评估,影响候选人的真实能力判断。本文提出首个西班牙语技能提取与分类数据集,设计标注方法以区分知识、技能与能力,并提供深度学习基线模型,推动该领域稳健解决方案的发展。

原文摘要 · Abstract (English)

Job application and assessment processes have evolved significantly in recent years, largely due to advancements in technology and changes in the way companies operate. Skill extraction and classification remain an important component of the modern hiring process as it provides a more objective way to evaluate candidates and automatically align their skills with the job requirements. However, to effectively evaluate the skills, the skill extraction tools must recognize varied mentions of skills on resumes, including direct mentions, implications, synonyms, acronyms, phrases, and proficiency levels, and differentiate between hard and soft skills. While tools like LLMs (Large Model Models) help extract and categorize skills from job applications, there's a lack of comprehensive datasets for evaluating the effectiveness of these models in accurately identifying and classifying skills in Spanish-language job applications. This gap hinders our ability to assess the reliability and precision of the models, which is crucial for ensuring that the selected candidates truly possess the required skills for the job. In this paper, we develop a Spanish language dataset for skill extraction and classification, provide annotation methodology to distinguish between knowledge, skill, and abilities, and provide deep learning baselines to advance robust solutions for skill classification.

技能分类西班牙语招聘系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。