arXiv:2509.20804cs.IR2025-09

Transformer模型在信息检索中表现不稳定,随机种子改变会导致结果波动大。

Performance Consistency of Learning Methods for Information Retrieval Tasks

  • 通过更换训练种子测试模型性能变化,评估稳定性。
  • 11个案例中9个F1得分标准差超0.075,7个精度标准差超0.125。
  • 提醒研究者需更严格评估方法可靠性,尤其针对深度学习模型。

为评估信息检索(IR)方法的性能准确性或鲁棒性,已有多种方法被提出。本文验证了使用测试集自助采样(bootstrapping)可有效估计性能波动。对于依赖初始种子的机器学习类方法,我们采用多组随机种子来考察性能变化。在三个不同IR任务中,测试了传统统计学习模型与基于Transformer的模型。结果显示,传统模型表现稳定,而Transformer模型随种子变化表现出显著波动:在11个案例中,9个F1分数(范围0.0–1.0)的标准差超过0.075;7个精确率(同范围)标准差超过0.125。值得注意的是,此前研究中低于0.02的差异即被视为方法改进证据,而当前波动远超此阈值。该发现揭示了Transformer模型对训练不稳定的敏感性,质疑了以往研究结果的可靠性,强调必须建立更严谨的评估流程。

原文摘要 · Abstract (English)

A range of approaches have been proposed for estimating the accuracy or robustness of the measured performance of IR methods. One is to use bootstrapping of test sets, which, as we confirm, provides an estimate of variation in performance. For IR methods that rely on a seed, such as those that involve machine learning, another approach is to use a random set of seeds to examine performance variation. Using three different IR tasks we have used such randomness to examine a range of traditional statistical learning models and transformer-based learning models. While the statistical models are stable, the transformer models show huge variation as seeds are changed. In 9 of 11 cases the F1-scores (in the range 0.0--1.0) had a standard deviation of over 0.075; while 7 of 11 precision values (also in the range 0.0--1.0) had a standard deviation of over 0.125. This is in a context where differences of less than 0.02 have been used as evidence of method improvement. Our findings highlight the vulnerability of transformer models to training instabilities and moreover raise questions about the reliability of previous results, thus underscoring the need for rigorous evaluation practices.

信息检索Transformer稳定性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。