验证查询性能预测方法组合的有效性,发现新方法仍可提升预测效果。
Combining Query Performance Predictors: A Reproducibility Study
- 测试多种查询性能预测方法组合,涵盖检索后与神经模型。
- 使用sMARE等新指标,结果支持早期组合策略有效性。
- 揭示不同方法间信息差异,为优化组合提供依据。
过去二十年来,众多查询性能预测(QPP)方法被提出。早在2009年,Hauff等人就探讨了不同QPP方法是否可通过组合提升预测质量。此后,相关方法与评估体系不断发展。本研究重新审视该工作,结合新型预测方法、评估指标与数据集,检验其结论的可复现性。研究扩展了先前工作:(i) 引入检索后方法,包括监督型神经网络(原研究仅限检索前方法);(ii) 采用sMARE指标,辅以传统相关系数与RMSE;(iii) 在Clueweb09B和TREC DL数据集上进行实验。结果总体支持原有结论,但也发现若干新现象。通过分析不同QPP方法间的相关性,进一步探究其是否捕捉互补信息或依赖重叠因素,从而深化对组合效果的理解。
原文摘要 · Abstract (English)
A large number of approaches to Query Performance Prediction (QPP) have been proposed over the last two decades. As early as 2009, Hauff et al. [28] explored whether different QPP methods may be combined to improve prediction quality. Since then, significant research has been done both on QPP approaches, as well as their evaluation. This study revisits Hauff et al.s work to assess the reproducibility of their findings in the light of new prediction methods, evaluation metrics, and datasets. We expand the scope of the earlier investigation by: (i) considering post-retrieval methods, including supervised neural techniques (only pre-retrieval techniques were studied in [28]); (ii) using sMARE for evaluation, in addition to the traditional correlation coefficients and RMSE; and (iii) experimenting with additional datasets (Clueweb09B and TREC DL). Our results largely support previous claims, but we also present several interesting findings. We interpret these findings by taking a more nuanced look at the correlation between QPP methods, examining whether they capture diverse information or rely on overlapping factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。