用小模型统一处理职位信息,提升理解精度与效率
Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
- 用合成数据微调小语言模型,兼顾分类与实体抽取
- 零样本泛化能力强,在结构化和非结构化场景表现优
- 多适配器架构降低维护成本,适合大规模工业应用
职位理解对领英连接人才与机会的使命至关重要。该任务需将非结构化、嘈杂的职位发布内容转化为标准化或衍生的职位属性,支撑多项领英产品。然而,构建可扩展、低成本且高性能的职位理解系统仍具挑战。本文提出一种由小型语言模型(SLM)驱动的统一语义建模框架。我们首先利用精心设计的合成任务(含推理轨迹)对开源SLM进行微调,这些任务同时针对基于分类体系的分类与无分类体系的实体抽取。这使模型在结构化与非结构化上下文中具备稳健的零样本泛化能力。在此基础上,引入按属性分组的多适配器架构,实现高效的任务特定适应,同时简化跨多种下游属性的模型管理。离线评估与在线A/B测试均显示性能显著提升,且运营复杂度降低。本工作为构建工业级文本理解系统提供了实用洞察。
原文摘要 · Abstract (English)
Job understanding is critical to LinkedIn's mission of connecting talent with opportunity. This task involves transforming unstructured and noisy job postings into standardized or derived job attributes that power numerous LinkedIn products. However, building a scalable, cost-efficient, and high-performing job understanding system remains challenging. In this paper, we present a unified semantic modeling framework powered by a small language model (SLM) to address the challenges. We begin by fine-tuning an open-source SLM using a suite of carefully curated synthetic tasks augmented with reasoning traces. These tasks jointly target taxonomy-guided classification and taxonomy-agnostic entity extraction. This allows the resulting model to acquire robust zero-shot generalization for job understanding in structured and unstructured contexts. Building upon this foundation, we introduce a multi-adapter architecture with attribute grouping to facilitate efficient task-specific adaptation while streamlining model management across diverse downstream attributes. Offline evaluations and online A/B tests demonstrate significant performance improvement while reducing operational complexity. Our work provides practical insights into building industry-scale text understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。