arXiv:2507.07924cs.IR2025-07被引 1

评估检索系统时,需同时关注真假阳性错误,才能准确判断标注数据质量。

Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems

  • 通过量化假阴性错误(Type II)补充传统假阳性分析
  • 发现平衡分类指标能更全面反映标注数据的判别能力
  • 适合关注评估可靠性与实验设计的检索研究者

信息检索(IR)系统的评估通常依赖带有人工标注相关性的查询-文档对(qrels),用于判断系统性能差异。由于人工标注成本高,高效标注方法被提出,需比较不同qrels的有效性。判别力(discriminative power)即正确识别系统间显著差异的能力至关重要。以往研究主要衡量被识别为显著不同的系统对比例及Ⅰ类错误(假阳性)。本文指出Ⅱ类错误(假阴性)同样重要,因其会导致错误科学结论。我们量化Ⅱ类错误,并提出使用平衡准确率等平衡分类指标来综合呈现qrels的判别力。实验基于多种替代标注方法生成的qrels进行验证,结果表明:结合Ⅱ类错误分析可获得更深入洞察,且平衡指标能以单一可比数值总结整体判别能力。

原文摘要 · Abstract (English)

The evaluation of Information Retrieval (IR) systems typically uses query-document pairs with corresponding human-labelled relevance assessments (qrels). These qrels are used to determine if one system is better than another based on average retrieval performance. Acquiring large volumes of human relevance assessments is expensive. Therefore, more efficient relevance assessment approaches have been proposed, necessitating comparisons between qrels to ascertain their efficacy. Discriminative power, i.e. the ability to correctly identify significant differences between systems, is important for drawing accurate conclusions on the robustness of qrels. Previous work has measured the proportion of pairs of systems that are identified as significantly different and has quantified Type I statistical errors. Type I errors lead to incorrect conclusions due to false positive significance tests. We argue that also identifying Type II errors (false negatives) is important as they lead science in the wrong direction. We quantify Type II errors and propose that balanced classification metrics, such as balanced accuracy, can be used to portray the discriminative power of qrels. We perform experiments using qrels generated using alternative relevance assessment methods to investigate measuring hypothesis testing errors in IR evaluation. We find that additional insights into the discriminative power of qrels can be gained by quantifying Type II errors, and that balanced classification metrics can be used to give an overall summary of discriminative power in one, easily comparable, number.

信息检索评估方法统计误差判别力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。