融合层级图结构的求职分类模型,提升招聘匹配精准度。
Hierarchical Job Classification with Similarity Graph Integration
- 构建共享嵌入空间,融合标准职业分类与自研层级体系
- 在大规模岗位数据集上显著优于现有方法
- 适合招聘系统优化与劳动力市场分析场景
在线招聘领域中,精准的职位分类对优化推荐系统、搜索排序和劳动力市场分析至关重要。随着职位标题与描述日益复杂,传统文本分类方法因无法充分挖掘行业类别的层次结构而表现受限。为此,本文提出一种新的表示学习与分类模型,将职位与层级行业类别共同嵌入到潜在空间中。模型整合了标准职业分类(SOC)系统与内部构建的层级分类体系Carotene,有效捕捉图结构与层次关系,从而提升分类准确率。通过共享嵌入空间,缓解冷启动问题,增强候选人与岗位的动态匹配能力。在大规模职位数据集上的实验表明,该模型在利用层次结构与丰富语义特征方面表现优异,显著超越现有方法。本研究为提升职位分类精度提供了稳健框架,助力招聘行业更科学决策。
原文摘要 · Abstract (English)
In the dynamic realm of online recruitment, accurate job classification is paramount for optimizing job recommendation systems, search rankings, and labor market analyses. As job markets evolve, the increasing complexity of job titles and descriptions necessitates sophisticated models that can effectively leverage intricate relationships within job data. Traditional text classification methods often fall short, particularly due to their inability to fully utilize the hierarchical nature of industry categories. To address these limitations, we propose a novel representation learning and classification model that embeds jobs and hierarchical industry categories into a latent embedding space. Our model integrates the Standard Occupational Classification (SOC) system and an in-house hierarchical taxonomy, Carotene, to capture both graph and hierarchical relationships, thereby improving classification accuracy. By embedding hierarchical industry categories into a shared latent space, we tackle cold start issues and enhance the dynamic matching of candidates to job opportunities. Extensive experimentation on a large-scale dataset of job postings demonstrates the model's superior ability to leverage hierarchical structures and rich semantic features, significantly outperforming existing methods. This research provides a robust framework for improving job classification accuracy, supporting more informed decision-making in the recruitment industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。