用一个小型语言模型统一工业搜索的查询理解,提升效果与效率。
A Unified Structured Query Understanding Framework for Industrial Semantic Search

- 用小语言模型统一处理多种查询理解任务,生成结构化结果。
- 在领英招聘系统中提升用户参与度,延迟低于100ms,GPU资源受限。
- 自动生成标注数据并评估模型,适合大规模工业搜索场景。
大规模工业搜索系统中的查询理解通常由多个独立、任务特定的组件构成。尽管各自可优化,但这种碎片化架构带来高维护成本和行为不一致问题,尤其对长尾查询表现不佳。本文提出并部署了一个统一的结构化查询理解系统,将异构功能整合至单一小型语言模型(SLM),实现基于模式约束的生成。为解决统一建模的数据瓶颈,引入Query Illuminator框架,兼具双重功能:(i) 作为教师模型实现高质量自动标注与知识蒸馏,(ii) 作为代理裁判在人工标签稀缺时实现可扩展评估。通过在领英招聘系统中的大量离线与在线测试验证该方法有效性。此外,还通过跨领域案例研究展示了系统在人员搜索中的横向扩展能力。结果显示,在满足严格低延迟服务要求(<100ms)及有限GPU资源条件下,显著提升用户参与度并降低运营成本。
原文摘要 · Abstract (English)
Query understanding in large-scale industrial search systems is typically implemented as a cascade of disparate, task-specific components. While individually optimizable, this fragmented architecture incurs high maintenance overhead and results in inconsistent behaviors, particularly for long-tail queries. In this work, we propose and deploy a unified structured query understanding system that consolidates these heterogeneous functions into a single Small Language Model (SLM) that performs schema-constrained generation. To address the data bottlenecks inherent in unified modeling, we introduce Query Illuminator, a dual-purpose framework serving as: (i) a teacher model for high-quality auto-annotation and distillation, and (ii) a surrogate judge for scalable evaluation where human labels are scarce. We validate this approach through extensive offline and online tests within LinkedIn's Job Search system. Furthermore, we demonstrate the framework's horizontal extensibility through a cross-domain case study on People Search. The results show improved user engagement and reduced operational costs, achieved while satisfying strict low-latency serving constraints on limited GPU resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。