用统一框架让一个模型搞定多种职场任务,性能还更高。
Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- 设计多任务排名基准WorkBench,整合六类职场任务
- 新模型UWE在未见任务上实现零样本排序,参数少100倍
- 适合需要低延迟、多任务的职场智能系统开发者
劳动市场智能应用需应对极端多标签目标空间、严格延迟约束及技能与职位名称等多模态文本。现有研究多为孤立的单任务模型。本文提出统一框架,构建首个涵盖六类职场任务的多任务排名基准WorkBench,基于真实本体与人工标注数据。跨任务分析显示显著正向迁移。据此提出通用职场嵌入(UWE),采用双向双编码器结构,通过多对多InfoNCE损失训练,并引入任务无关的软后期交互机制。UWE在未见目标空间上实现零样本排名,参数量仅为最优通用模型Qwen3-8B的百分之一,延迟降低两个数量级,且提升4.4 MAP。
原文摘要 · Abstract (English)
Applications in labor market intelligence demand specialized NLP systems for a wide range of tasks, characterized by extreme multi-label target spaces, strict latency constraints, and multiple text modalities such as skills and job titles. These constraints have led to isolated, task-specific developments in the field, with models and benchmarks focused on single prediction tasks. Exploiting the shared structure of work-related data, we propose a unifying framework, combining a wide range of tasks in a multi-task ranking benchmark, and a flexible architecture tackling text-driven work tasks with a single model. The benchmark, WorkBench, is the first unified evaluation suite spanning six work-related tasks formulated explicitly as ranking problems, curated from real-world ontologies and human-annotated resources. WorkBench enables cross-task analysis, where we find significant positive cross-task transfer. This insight leads to Unified Work Embeddings (UWE), a task-agnostic bi-encoder that exploits our training-data structure with a many-to-many InfoNCE objective, and leverages token-level embeddings with task-agnostic soft late interaction. UWE demonstrates zero-shot ranking performance on unseen target spaces in the work domain, and enables low-latency inference with two orders of magnitude fewer parameters than best-performing generalist models (Qwen3-8B), with +4.4 MAP improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。