arXiv:2503.08965cs.IR2025-03被引 6

用大模型生成搜索结果有用性标签,提升评估效率与准确性。

LLM-Driven Usefulness Labeling for IR Evaluation

  • 结合查询、文档及用户行为信号,让大模型理解完整搜索任务。
  • 在短会话中,提供上下文的大模型能更准确判断结果有用性。
  • 验证了相关性与有用性的差异,适合评估系统优化者参考。

在信息检索领域,评估对优化搜索体验和满足多样用户意图至关重要。近年来,随着大模型的发展,研究开始探索自动化文档相关性标签的生成,以替代传统依赖众包人工标注的方式——该方式耗时且成本高。本研究聚焦于大模型生成的有用性标签,这一关键评估指标能更好反映用户搜索意图与任务目标,是传统相关性评估的不足之处。实验综合使用任务级、查询级、文档级特征以及用户搜索行为信号,这些因素对定义文档有用性至关重要。研究发现:(i)预训练大模型可通过理解完整的搜索会话生成中等水平的有用性标签;(ii)在短搜索会话中,提供会话上下文的大模型判断表现更优。此外,我们还考察了大模型是否能捕捉相关性与有用性之间的差异,并通过消融实验识别出生成准确有用性标签最关键的指标。结论表明,本研究通过评估关键指标并优化实用性,探索了大模型生成有用性标签的有效路径。

原文摘要 · Abstract (English)

In the information retrieval (IR) domain, evaluation plays a crucial role in optimizing search experiences and supporting diverse user intents. In the recent LLM era, research has been conducted to automate document relevance labels, as these labels have traditionally been assigned by crowd-sourced workers - a process that is both time and consuming and costly. This study focuses on LLM-generated usefulness labels, a crucial evaluation metric that considers the user's search intents and task objectives, an aspect where relevance falls short. Our experiment utilizes task-level, query-level, and document-level features along with user search behavior signals, which are essential in defining the usefulness of a document. Our research finds that (i) pre-trained LLMs can generate moderate usefulness labels by understanding the comprehensive search task session, (ii) pre-trained LLMs perform better judgement in short search sessions when provided with search session contexts. Additionally, we investigated whether LLMs can capture the unique divergence between relevance and usefulness, along with conducting an ablation study to identify the most critical metrics for accurate usefulness label generation. In conclusion, this work explores LLM-generated usefulness labels by evaluating critical metrics and optimizing for practicality in real-world settings.

信息检索大模型评估有用性标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。