用大模型提升Pinterest搜索相关性,效果显著。
Improving Pinterest Search Relevance Using Large Language Models
- 融合大模型生成的图文描述与多种文本数据构建相关性模型
- 半监督学习扩展训练数据,实现多语言支持
- 模型压缩后部署,兼顾效果与实时性能
为提升Pinterest搜索的相关性评分,我们将在搜索相关性模型中集成大型语言模型(LLMs),利用精心设计的文本表示来有效预测贴文(Pins)的相关性。方法结合搜索查询与内容表征,包括由生成式视觉语言模型提取的标题、链接文本数据、历史高质量互动查询、用户收藏的画板、贴文标题和描述,构建鲁棒的相关性预测模型。采用半监督学习策略,高效扩充训练数据,突破昂贵的人工标注数据限制。通过使用多语言LLMs,系统可覆盖未见语言和领域,尽管初始数据和标注者专长仅限于英语。此外,将基于LLM的模型提炼为可实时服务的模型架构与特征。我们提供了全面的离线实验验证,并在大规模部署后展示显著提升的效果。
原文摘要 · Abstract (English)
To improve relevance scoring on Pinterest Search, we integrate Large Language Models (LLMs) into our search relevance model, leveraging carefully designed text representations to predict the relevance of Pins effectively. Our approach uses search queries alongside content representations that include captions extracted from a generative visual language model. These are further enriched with link-based text data, historically high-quality engaged queries, user-curated boards, Pin titles and Pin descriptions, creating robust models for predicting search relevance. We use a semi-supervised learning approach to efficiently scale up the amount of training data, expanding beyond the expensive human labeled data available. By utilizing multilingual LLMs, our system extends training data to include unseen languages and domains, despite initial data and annotator expertise being confined to English. Furthermore, we distill from the LLM-based model into real-time servable model architectures and features. We provide comprehensive offline experimental validation for our proposed techniques and demonstrate the gains achieved through the final deployed system at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。