arXiv:2509.15786cs.AIcs.IR2025-09EMNLP被引 4

自动构建更准确、可扩展的职业分类体系。

Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration

  • 基于语义聚类与多智能体协作,从原始招聘信息中自动生成职业分类。
  • 在三个真实数据集上,分类体系一致性与可扩展性优于现有方法。
  • 适合关注职业分析、劳动力市场研究的从业者和研究人员。

构建稳健的职业分类体系对职位推荐、劳动力市场洞察等应用至关重要。人工梳理耗时费力,而现有自动化方法或难以适应动态区域市场(自上而下),或在噪声数据中难以形成连贯层级(自下而上)。我们提出CLIMB(基于聚类的多智能体分类构建框架),完全自动化地从原始职位发布信息中生成高质量、数据驱动的职业分类。CLIMB首先通过全局语义聚类提炼核心职业,再利用基于反思的多智能体系统迭代构建连贯层级结构。在三个不同且真实的大型数据集上,实验表明CLIMB生成的分类体系在一致性与可扩展性方面均优于现有方法,并能有效捕捉区域特性。代码与数据已开源:https://anonymous.4open.science/r/CLIMB。

原文摘要 · Abstract (English)

Creating robust occupation taxonomies, vital for applications ranging from job recommendation to labor market intelligence, is challenging. Manual curation is slow, while existing automated methods are either not adaptive to dynamic regional markets (top-down) or struggle to build coherent hierarchies from noisy data (bottom-up). We introduce CLIMB (CLusterIng-based Multi-agent taxonomy Builder), a framework that fully automates the creation of high-quality, data-driven taxonomies from raw job postings. CLIMB uses global semantic clustering to distill core occupations, then employs a reflection-based multi-agent system to iteratively build a coherent hierarchy. On three diverse, real-world datasets, we show that CLIMB produces taxonomies that are more coherent and scalable than existing methods and successfully capture unique regional characteristics. We release our code and datasets at https://anonymous.4open.science/r/CLIMB.

职业分类语义聚类多智能体数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。