用微调大模型自动评估搜索相关性,效率提升且结果可靠。
LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
- 用微调LLM替代人工标注进行搜索相关性评估
- 模型判断与人工标注高度一致,显著降低实验最小可检测效应
- 适合大规模在线实验的快速迭代与多场景评估
相关性评估在个性化搜索系统中至关重要,确保搜索结果与用户查询意图匹配。传统依赖人工标注的方法成本高、周期长,难以扩展。本文介绍Pinterest搜索团队采用微调大模型自动化相关性评估的方法。通过严格验证,证明大模型生成的判断与人工标注高度一致,可在保证可靠性的同时大幅提升评估效率。基于大模型标注,还可拓展查询集、优化采样设计,实现对更广泛搜索体验的大规模高效评估。该方法显著提升了相关性度量质量,并大幅降低在线实验中的最小可检测效应(MDE)。
原文摘要 · Abstract (English)
Relevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user's queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present our approach at Pinterest Search to automate relevance evaluation for online experiments using fine-tuned LLMs. We rigorously validate the alignment between LLM-generated judgments and human annotations, demonstrating that LLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging LLM-based labeling further unlocks the opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effect (MDE) in online experiment measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。