arXiv:2507.09488cs.IR2025-07中稿 · ICTIR 2025被引 11

用多维度标准提升LLM判断相关性的准确与可解释性

Criteria-Based LLM Relevance Judgments

  • 将相关性拆解为精确性、覆盖度等多标准,避免单一评分偏差
  • 在TREC DL 2019/2020及LLMJudge数据集上提升排名表现
  • 适合需要高可信自动评估的检索系统研发者参考

相关性判断对信息检索系统评估至关重要,但传统人工标注耗时费力。许多研究转向自动替代方案,其中大语言模型(LLM)通过提示词直接生成相关性标签具有可扩展性。然而,无约束提示常导致错误预测且输出难以理解。本文提出多准则框架,将相关性分解为精确性、覆盖度、主题相关性和上下文契合度等多个维度,相比直接评分法显著提升评估的鲁棒性与可解释性。我们在TREC深度学习2019、2020年赛道及基于TREC DL 2023的LLMJudge数据集上验证该方法,结果表明多准则判断能有效提升系统排名表现。同时分析了其相对于直接评分的优势与局限,为未来自动评估框架设计提供关键洞见。

原文摘要 · Abstract (English)

Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers turn to automatic alternatives to accelerate method development. Among these, Large Language Models (LLMs) provide a scalable solution by generating relevance labels directly through prompting. However, prompting an LLM for a relevance label without constraints often results in not only incorrect predictions but also outputs that are difficult for humans to interpret. We propose the Multi-Criteria framework for LLM-based relevance judgments, decomposing the notion of relevance into multiple criteria--such as exactness, coverage, topicality, and contextual fit--to improve the robustness and interpretability of retrieval evaluations compared to direct grading methods. We validate this approach on three datasets: the TREC Deep Learning tracks from 2019 and 2020, as well as LLMJudge (based on TREC DL 2023). Our results demonstrate that Multi-Criteria judgments enhance the system ranking/leaderboard performance. Moreover, we highlight the strengths and limitations of this approach relative to direct grading approaches, offering insights that can guide the development of future automatic evaluation frameworks in information retrieval.

信息检索LLM评估多准则判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。