提出更可靠的多系统检索效果比较方法,提升实验结果可信度。
Towards Reliable Testing for Multiple Information Retrieval System Comparisons
- 采用威尔科克斯检验结合贝尼杰曼校正处理多系统对比
- 在典型样本量下保持显著性水平的类型Ⅰ错误率
- 兼具高统计功效,适合真实检索系统评估场景
假设检验是评估信息检索系统性能差异的主流工具。研究人员通过统计检验判断实验室观察到的差异是否能在实际线上环境推广,而非仅由样本偶然导致。已有大量研究关注成对系统比较中哪种检验最可靠,但多数真实实验涉及超过两个系统。在多重比较场景下,同时测试多个系统可能放大检验错误率。本文通过模拟和真实TREC数据,评估多种多重比较方法的可靠性。实验表明,在典型样本量下,威尔科克斯检验配合贝尼杰曼校正能准确控制类型Ⅰ错误率,且在统计功效上表现最优。
原文摘要 · Abstract (English)
Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests to check whether those differences will generalise to online settings or are just due to the samples observed in the laboratory. Much work has been devoted to studying which test is the most reliable when comparing a pair of systems, but most of the IR real-world experiments involve more than two. In the multiple comparisons scenario, testing several systems simultaneously may inflate the errors committed by the tests. In this paper, we use a new approach to assess the reliability of multiple comparison procedures using simulated and real TREC data. Experiments show that Wilcoxon plus the Benjamini-Hochberg correction yields Type I error rates according to the significance level for typical sample sizes while being the best test in terms of statistical power.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。