arXiv:2604.25349cs.IRstat.AP2026-04

警告:信息检索中滥用威尔科克斯检验,反而导致错误结论。

Stop Using the Wilcoxon Test: Myth, Misconception and Misuse in IR Research

  • 用系统性分析揭示威尔科克斯检验在信息检索中的根本缺陷
  • 实验证明其第一类错误率失控,结果不可靠
  • 适合关注评估方法严谨性的研究人员参考

在信息检索系统评测中,威尔科克斯符号秩检验常被视为比t检验更安全的非参数替代方法,这一观点源于教材和建议中将该检验视为因指标分数不满足正态分布而应使用的正确选择。我们指出这一叙事具有误导性且有害。对统计学教材的细致审查显示,相关假设的表述存在矛盾与遗漏,导致研究者长期误解。实际上,威尔科克斯检验在信息检索场景下极易失控,第一类错误率严重偏离预期水平,产生虚假的安全感,同时引入更严重的偏差风险,几乎必然导致测试失效并误导研究结论。通过文献综述、理论分析及TREC数据的实证演示,我们阐明了其失灵机制。因此,继续在信息检索评估中使用威尔科克斯检验缺乏正当性,放弃该方法将显著提升本领域的方法论严谨性。

原文摘要 · Abstract (English)

In benchmarking of Information Retrieval systems, the Wilcoxon signed-rank test is often treated as a safer alternative to the t-test. This belief is fueled by textbooks and recommendations that portray Wilcoxon as the proper non-parametric alternative because metric scores are not normally distributed. We argue that this narrative is misleading and harmful. A careful review of Statistics textbooks reveals inconsistencies and omissions in how the assumptions underlying these tests are presented, fostering confusion that has propagated into IR research. As a result, Wilcoxon has been routinely misapplied for decades, creating a false sense of safety against a threat that was never there to begin with, while introducing another one so severe that it virtually guarantees the test will break down and mislead researchers. Through a combination of systematic literature review, analysis and empirical demonstrations with TREC data, we show how and why the Wilcoxon test easily loses control of its Type I error rate in IR settings. We conclude that the continued use of Wilcoxon in IR evaluation is unjustified and that abandoning it would improve the methodological soundness of our field.

统计检验信息检索方法论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。