arXiv:2504.01101cs.IRcs.LG2025-04被引 4

现有查询性能预测方法在跨数据集时表现不稳定,难以有效指导查询处理优化。

Uncovering the Limitations of Query Performance Prediction: Failures, Insights, and Implications for Selective Query Processing

  • 对比多种稀疏与稠密检索器,评估主流性能预测方法的泛化能力
  • 在不同数据集上预测准确率差异显著,最高相差超20个百分点
  • 适用于特定场景但缺乏通用性,对智能查询处理帮助有限

查询性能预测(QPP)可估计检索系统对特定查询的有效性,为搜索效果和查询处理提供重要参考。尽管研究广泛,现有方法在跨不同检索范式和数据集时仍面临严重泛化挑战。本文全面评估了最新QPP方法(如NQC、UQC)、LETOR特征及新提出的稠密预测器,使用多种稀疏排序器(BM25、DFree及带查询扩展版本)与混合或稠密排序器(SPLADE、ColBERT),在ROBUST、GOV2、WT10G和MS MARCO等多个测试集上分析预测值与实际性能的关系,重点关注泛化性与鲁棒性。结果表明,预测准确率存在显著波动,数据集是主要影响因素,其次为排序器类型。部分稀疏预测器在TREC ROBUST和GOV2上表现尚可,但在WT10G和MS-MARCO上失效。虽某些预测器在特定场景下有潜力,但整体局限性限制其应用价值。实验显示,基于QPP的选查询处理仅带来微弱增益,凸显出需开发能跨数据集泛化、适配稠密检索架构且实用的新预测方法。

原文摘要 · Abstract (English)

Query Performance Prediction (QPP) estimates retrieval systems effectiveness for a given query, offering valuable insights for search effectiveness and query processing. Despite extensive research, QPPs face critical challenges in generalizing across diverse retrieval paradigms and collections. This paper provides a comprehensive evaluation of state-of-the-art QPPs (e.g. NQC, UQC), LETOR-based features, and newly explored dense-based predictors. Using diverse sparse rankers (BM25, DFree without and with query expansion) and hybrid or dense (SPLADE and ColBert) rankers and diverse test collections ROBUST, GOV2, WT10G, and MS MARCO; we investigate the relationships between predicted and actual performance, with a focus on generalization and robustness. Results show significant variability in predictors accuracy, with collections as the main factor and rankers next. Some sparse predictors perform somehow on some collections (TREC ROBUST and GOV2) but do not generalise to other collections (WT10G and MS-MARCO). While some predictors show promise in specific scenarios, their overall limitations constrain their utility for applications. We show that QPP-driven selective query processing offers only marginal gains, emphasizing the need for improved predictors that generalize across collections, align with dense retrieval architectures and are useful for downstream applications.

查询预测检索系统泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。