arXiv:2503.12989cs.CLcs.AI2025-03AAAI被引 2

用分阶段框架提升大模型职业分类能力,降低成本。

A Multi-Stage Framework with Taxonomy-Guided Reasoning for Occupation Classification Using Large Language Models

  • 分三阶段:推理、检索、重排序,融入职业分类知识
  • 在大规模数据集上性能优于小模型,接近GPT-4o水平
  • 适合需低成本高效率职业分类的机构或研究者

自动将工作数据标注为标准化职业类别,即职业分类,对劳动力市场分析至关重要。然而,该任务常受制于数据稀缺和人工标注难题。尽管大语言模型(LLMs)凭借其广泛世界知识和上下文学习能力具有潜力,但其在职业分类体系上的知识掌握程度尚不明确。本研究评估了LLMs从分类体系中生成精确职业实体的能力,揭示了小模型的局限性。为此,我们提出一种多阶段框架,包含推理、检索和重排序阶段,通过引入基于分类体系的推理示例,使输出更契合分类知识。在大规模数据集上的评估表明,该框架不仅提升了职业与技能分类性能,还提供了一种比前沿模型如GPT-4o更具成本效益的替代方案,在显著降低计算开销的同时保持强性能。这使其成为各类大模型上职业分类及相关任务的实用且可扩展的解决方案。

原文摘要 · Abstract (English)

Automatically annotating job data with standardized occupations from taxonomies, known as occupation classification, is crucial for labor market analysis. However, this task is often hindered by data scarcity and the challenges of manual annotations. While large language models (LLMs) hold promise due to their extensive world knowledge and in-context learning capabilities, their effectiveness depends on their knowledge of occupational taxonomies, which remains unclear. In this study, we assess the ability of LLMs to generate precise taxonomic entities from taxonomy, highlighting their limitations, especially for smaller models. To address these challenges, we propose a multi-stage framework consisting of inference, retrieval, and reranking stages, which integrates taxonomy-guided reasoning examples to enhance performance by aligning outputs with taxonomic knowledge. Evaluations on a large-scale dataset show that our framework not only enhances occupation and skill classification tasks, but also provides a cost-effective alternative to frontier models like GPT-4o, significantly reducing computational costs while maintaining strong performance. This makes it a practical and scalable solution for occupation classification and related tasks across LLMs.

职业分类大模型多阶段框架知识对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。