PinCLIP提升图片与文本对齐,显著改善Pinterest推荐效果。
PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest
- 用混合视觉变压器融合图文信息,捕捉多粒度特征。
- 在检索任务中比Qwen提升20%,在线测试带动用户参与度增长。
- 有效缓解冷启动问题,新内容分享率提高15%。
尽管多模态视觉语言模型(VLMs)在多个领域表现优异,但在推荐和检索系统中的应用仍面临训练目标不一致和服务效率瓶颈等问题。本文提出PinCLIP,一种大规模视觉表征学习方法,旨在通过利用VLM实现图像与文本对齐,提升Pinterest的检索与排序模型性能。我们设计了一种新型混合视觉变压器架构,采用VLM主干网络和混合融合机制,在不同粒度上捕捉多模态内容表示。除标准的图像-文本对齐目标外,还引入邻居对齐目标,建模Pinterest Pin-Board图中多模态表示间的交叉融合。离线评估显示,PinCLIP在多模态检索任务中比SOTA基线(如Qwen)高出20%;在线A/B测试表明,各主要页面均带来显著业务提升。特别地,其显著缓解了冷启动问题,使新内容的有机转贴率提升15%,新广告点击率提高8.7%。
原文摘要 · Abstract (English)
While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems remains a challenge, due to issues like training objective discrepancies and serving efficiency bottlenecks. This paper introduces PinCLIP, a large-scale visual representation learning approach developed to enhance retrieval and ranking models at Pinterest by leveraging VLMs to learn image-text alignment. We propose a novel hybrid Vision Transformer architecture that utilizes a VLM backbone and a hybrid fusion mechanism to capture multi-modality content representation at varying granularities. Beyond standard image-to-text alignment objectives, we introduce a neighbor alignment objective to model the cross-fusion of multi-modal representations within the Pinterest Pin-Board graph. Offline evaluations show that PinCLIP outperforms state-of-the-art baselines, such as Qwen, by 20% in multi-modal retrieval tasks. Online A/B testing demonstrates significant business impact, including substantial engagement gains across all major surfaces in Pinterest. Notably, PinCLIP significantly addresses the "cold-start" problem, enhancing fresh content distribution with a 15% Repin increase in organic content and 8.7% higher click for new Ads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。