arXiv:2505.24640cs.CLcs.AI2025-05被引 6

用轻量模型高效提取岗位技能,提升分析速度与精度。

Efficient Text Encoders for Labor Market Analysis

  • 采用对比学习与词粒度注意力机制,优化技能分类效率。
  • 在新基准上达到顶尖性能,轻量模型推理速度快。
  • 适合需要实时处理海量岗位数据的机构或平台使用。

劳动力市场分析依赖于从岗位广告中提取信息,这些信息虽有价值但呈非结构化状态,包含职位名称和技能要求。现有先进技能提取方法虽表现优异,但依赖大型语言模型(LLMs),计算成本高且响应慢。本文提出 extbf{ConTeXT-match},一种适用于极端多标签分类任务的新型对比学习方法,结合词粒度注意力机制,显著提升技能提取的效率与性能,使用轻量级双编码器模型即达当前最优结果。为支持稳健评估,我们引入 extbf{Skill-XL} 基准,提供详尽的句子级技能标注,明确解决大规模标签空间中的冗余问题。此外,我们还推出 extbf{JobBERT V2},一个改进的职位名称规范化模型,利用提取的技能生成高质量职位表示。实验表明,所提模型高效、准确且可扩展,适用于大规模、实时的劳动力市场分析。

原文摘要 · Abstract (English)

Labor market analysis relies on extracting insights from job advertisements, which provide valuable yet unstructured information on job titles and corresponding skill requirements. While state-of-the-art methods for skill extraction achieve strong performance, they depend on large language models (LLMs), which are computationally expensive and slow. In this paper, we propose \textbf{ConTeXT-match}, a novel contrastive learning approach with token-level attention that is well-suited for the extreme multi-label classification task of skill classification. \textbf{ConTeXT-match} significantly improves skill extraction efficiency and performance, achieving state-of-the-art results with a lightweight bi-encoder model. To support robust evaluation, we introduce \textbf{Skill-XL}, a new benchmark with exhaustive, sentence-level skill annotations that explicitly address the redundancy in the large label space. Finally, we present \textbf{JobBERT V2}, an improved job title normalization model that leverages extracted skills to produce high-quality job title representations. Experiments demonstrate that our models are efficient, accurate, and scalable, making them ideal for large-scale, real-time labor market analysis.

技能提取轻量模型劳动市场分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。