arXiv:2509.25602cs.IR2025-09被引 9

提出可复现的LLM相关性判断框架,提升检索评估可靠性

TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval

  • 基于任务感知的评分标准,迭代优化判断流程
  • 在TREC DL 2019/2020和LLMJudge上表现优异
  • 适合需标准化评估的检索系统研究者

基于大语言模型的相关性判断已成为信息检索评估的重要方法,虽在LLMJudge排行榜上展现出与人工判断高度相关的性能,但现有方法严重依赖敏感提示策略,缺乏标准化生成流程。为此,我们重新引入任务感知的评分标准评估框架(TRUE),该框架原用于搜索会话中的有用性评估,现扩展至相关性判断,以解决其因工作流可复现性强而表现出的有效性缺口。该框架通过迭代数据采样与推理,综合评估意图、覆盖度、具体性、准确性和有用性等多个维度。我们在TREC DL 2019、2020及LLMJudge数据集上验证了TRUE,结果表明其在系统排名类LLM排行榜上表现强劲。本工作核心在于提供一个可复现的基于LLM的相关性判断框架,并进一步分析其多维度有效性。

原文摘要 · Abstract (English)

LLM-based relevance judgment generation has become a crucial approach in advancing evaluation methodologies in Information Retrieval (IR). It has progressed significantly, often showing high correlation with human judgments as reflected in LLMJudge leaderboards \cite{rahmani2025judging}. However, existing methods for relevance judgments, rely heavily on sensitive prompting strategies, lacking standardized workflows for generating reliable labels. To fill this gap, we reintroduce our method, \textit{Task-aware Rubric-based Evaluation} (TRUE), for relevance judgment generation. Originally developed for usefulness evaluation in search sessions, we extend TRUE to mitigate the gap in relevance judgment due to its demonstrated effectiveness and reproducible workflow. This framework leverages iterative data sampling and reasoning to evaluate relevance judgments across multiple factors including intent, coverage, specificity, accuracy and usefulness. In this paper, we evaluate TRUE on the TREC DL 2019, 2020 and LLMJudge datasets and our results show that TRUE achieves strong performance on the system-ranking LLM leaderboards. The primary focus of this work is to provide a reproducible framework for LLM-based relevance judgments, and we further analyze the effectiveness of TRUE across multiple dimensions.

信息检索大模型评估可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。