用用户历史行为数据提升大模型对搜索结果的评估准确性。
As It Was: Aligning LLM Search Evaluation with Historical User Preferences

- 为每个搜索结果添加历史点击数据卡片,结合行为证据做判断。
- 在音乐搜索中相关性提升5%,分歧案例改善91%。
- 适合需要真实用户反馈支持的搜索系统评估场景。
大规模搜索系统演进速度远超人工质量保障能力,尤其针对长尾意图和多语言查询。基于大模型的评判方法虽可扩展,但仅依赖语义相似性或世界知识会偏离真实用户偏好,尤其在模糊查询下。本文提出一种行为基础的大模型评判器,为每个搜索结果页面(SERP)项附加轻量级、可审计的行为先验——查询-相关性-曝光(QRI)卡片,汇总历史用户对相似查询与结果的互动情况,提供紧凑的实证依据,帮助模型在保持语义推理的同时解决歧义,提升判断一致性。在Spotify大规模音乐搜索评估中,基于6000个重构的SERP的历史交互数据,该方法使相关性与用户偏好对齐度整体提升约5%,分歧案例改善达91%。在跨五种语言的人工标注数据集上,进一步提升与人类评判的相关性15%。更重要的是,在与线上A/B测试结果对比时,该方法始终表现出更高的一致性。尽管绝对对齐度仍中等,但结果表明轻量行为引导能显著增强大模型评估在真实搜索系统中的可靠性与实用性。
原文摘要 · Abstract (English)
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or world knowledge can drift from actual user preferences, particularly for ambiguous queries. We introduce a behavior-grounded LLM judge that augments each SERP item with a lightweight and auditable behavioral prior in the form of a Query-Relevance-Impressions (QRI) card. Each card summarizes how users have historically interacted with similar queries and results, providing compact empirical evidence that the judge can cite to resolve ambiguity and make more consistent relevance judgments while still relying on semantic reasoning. In a large-scale music search evaluation at Spotify, using relevance estimates derived from historical user interactions across 6,000 recomposed SERPs, the behavior-grounded judge achieves stronger alignment with user preferences, improving Spearman rank correlation by approximately 5% overall and yielding a 91% relative improvement on disagreement cases. On a multilingual human-judged dataset spanning five languages, grounding further increases correlation with human relevance judgments by 15%. Importantly, when evaluated against outcomes from a live A/B test, the grounded judge shows consistently higher alignment with the observed winning model. While absolute alignment remains moderate, these findings demonstrate that lightweight behavioral grounding can improve the reliability and practical usefulness of LLM-based evaluation in real-world search systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。