arXiv:2608.02446cs.IRcs.LG2026-08

用视觉语言模型自动评估搜索相关性,提升效率与精度。

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

论文配图:Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
图 1 · 摘自论文原文
  • 用视觉语言模型替代人工标注,实现自动化相关性判断。
  • 模型判断与人工标注高度一致,显著降低实验最小可检测效应。
  • 适合大规模搜索系统优化,尤其在需要快速迭代的场景中

相关性评估在个性化搜索系统中至关重要,是确保搜索结果与用户查询意图对齐的关键指标。传统依赖人工标注的方法成本高、周期长,难以扩展。本文提出一种基于视觉语言模型(VLM)的自动化相关性评估流程,已在Pinterest搜索系统中部署并用于在线A/B实验。我们严格验证了VLM生成判断与人工标注的一致性,证明VLM能提供可靠的相关性度量,极大提升评估效率。借助VLM标注,还可拓展查询集、优化采样设计,高效评估更大范围的搜索体验。该方法提升了相关性度量质量,并显著降低了在线实验中的最小可检测效应(Minimum Detectable Effects, MDE)。

原文摘要 · Abstract (English)

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.

搜索相关性视觉语言模型自动化评估在线实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。