arXiv:2604.04140cs.IR2026-04被引 3

用结构化信息需求提升大模型评估可靠性

Formalized Information Needs Improve Large-Language-Model Relevance Judgments

  • 用大模型自动生成符合标准的检索主题
  • 无结构化主题时,大模型判断相关文档更多且一致性更低
  • 即使主题不完全匹配人类,也能提升人机评估一致率

基于Cranfield的检索评估中,相关文档数量过少或过多,以及人工评估者间相关性低,都会降低评估结果的可靠性。传统人工评估常将信息需求形式化为包含描述和叙述的检索主题,以控制相关文档数量并保持高一致性。然而,当前使用大语言模型(LLM)作为评估者的新型评测设置通常仅使用查询,可能导致评估可靠性下降。为此,我们利用大模型将信息需求合成生成符合先前人工评估标准的检索主题(含描述与叙述)。在TREC Deep Learning和Robust04 2019/2020版数据集上,对比了使用合成主题的评估者与默认仅用查询的评估者。结果表明,未形式化的评估者判定的相关文档显著更多,且评估一致性更低,导致检索评估可靠性下降。此外,即使合成主题与人工主题相似度不高,其仍能提升人类与大模型之间相关性判断的一致性。研究结果表明,大模型评估应像人工评估一样采用形式化信息需求,当缺乏人工形式化时,可借助大模型自动生成主题,以提高评估可靠性。

原文摘要 · Abstract (English)

Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability.

大模型评估信息检索主题形式化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。