用摘要代替全文评估检索系统,效率提升但可能引入偏差。
The Effect of Document Summarization on LLM-Based Relevance Judgments
- 用LLM生成文档摘要替代原文输入,评估模型判断
- 摘要判断与完整文本判断在排名稳定性上相当
- 不同模型和数据集下存在系统性标签偏差
相关性判断是信息检索(IR)系统评估的核心,但由人工标注成本高、耗时长。大语言模型(LLMs)近年来被提出作为自动评估工具,表现出与人工标注的良好一致性。以往研究通常将文档视为固定单元,直接输入全文内容给LLM评估。本文研究文本摘要对基于LLM的判断可靠性及其在下游IR评估中的影响。我们在多个TREC数据集上,使用先进LLM比较了基于全文与基于不同长度的LLM生成摘要的判断结果。考察其与人工标签的一致性、对检索效果评估的影响以及对系统排序稳定性的作用。结果表明,基于摘要的判断在系统排名稳定性上可媲美全文判断,但引入了系统性标签分布偏移,且偏差随模型和数据集而异。这些发现表明,摘要化评估既为大规模高效评估提供了机会,也是一种具有重要影响的方法选择。
原文摘要 · Abstract (English)
Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as automated assessors, showing promising alignment with human annotations. Most prior studies have treated documents as fixed units, feeding their full content directly to LLM assessors. We investigate how text summarization affects the reliability of LLM-based judgments and their downstream impact on IR evaluation. Using state-of-the-art LLMs across multiple TREC collections, we compare judgments made from full documents with those based on LLM-generated summaries of different lengths. We examine their agreement with human labels, their effect on retrieval effectiveness evaluation, and their influence on IR systems' ranking stability. Our findings show that summary-based judgments achieve comparable stability in systems' ranking to full-document judgments, while introducing systematic shifts in label distributions and biases that vary by model and dataset. These results highlight summarization as both an opportunity for more efficient large-scale IR evaluation and a methodological choice with important implications for the reliability of automatic judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。