用大模型生成网页搜索有用性标签,提升评估准确性。
LLM-Driven Usefulness Judgment for Web Search Evaluation
- 基于任务的评分框架,结合搜索会话上下文迭代推理。
- 微调后的大模型在结构化上下文中判断更有用性,效果更优。
- 适合需要评估用户真实目标达成度的搜索系统优化者。
评估是优化搜索体验、支持多样化用户意图的关键。传统方法依赖相关性标签,但仅靠相关性无法衡量系统帮助用户达成目标的能力,因此有用性成为重要评估指标。本文提出基于大模型的有用性评估方法——任务感知的评分框架(TRUE),通过迭代采样与推理建模复杂搜索行为。研究发现:(i) 大模型可利用包含个性化和上下文理解的完整搜索会话历史,生成中等水平的有用性标签;(ii) 在提供结构化搜索会话上下文时,微调后的大模型显著提升判断质量。此外,我们检验了大模型区分相关性与有用性的能力,尤其在二者分歧影响搜索成功时。通过消融实验识别关键指标,兼顾生成效率与成本效益。本研究推动了基于大模型的有用性评估发展,优化关键用户指标,验证标签可靠性,并确保大规模部署可行性。
原文摘要 · Abstract (English)
Evaluation is fundamental in optimizing search experiences and supporting diverse user intents in Information Retrieval (IR). Traditional search evaluation methods primarily rely on relevance labels, which assess how well retrieved documents match a user's query. However, relevance alone fails to capture a search system's effectiveness in helping users achieve their search goals, making usefulness a critical evaluation criterion. In this paper, we explore an alternative approach: LLM-generated usefulness labels, which incorporate both implicit and explicit user behavior signals to evaluate document usefulness. We propose Task-aware Rubric-based Usefulness Evaluation (TRUE), a rubric-driven evaluation method that employs iterative sampling and reasoning to model complex search behavior patterns. Our findings show that (i) LLMs can generate moderate usefulness labels by leveraging comprehensive search session history incorporating personalization and contextual understanding, and (ii) fine-tuned LLMs improve usefulness judgments when provided with structured search session contexts. Additionally, we examine whether LLMs can distinguish between relevance and usefulness, particularly in cases where this divergence impacts search success. We also conduct an ablation study to identify key metrics for accurate usefulness label generation, optimizing for token efficiency and cost-effectiveness in real-world applications. This study advances LLM-based usefulness evaluation by refining key user metrics, exploring LLM-generated label reliability, and ensuring feasibility for large-scale search systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。