arXiv:2601.17359cs.IR2026-01被引 3

提出三种新评估框架,更全面衡量查询性能预测效果

Breaking Flat: A Generalised Query Performance Prediction Evaluation Framework

  • 构建三类评估场景,涵盖单/多查询与单/多排序器组合
  • 发现预测最佳排序器比预测查询难易度更困难
  • 揭示不同任务下模型表现差异大,需针对性评估

传统查询性能预测(QPP)旨在识别给定排序模型下哪些查询表现好、哪些差。更精细且更具挑战性的扩展是确定对特定查询最有效的排序模型。本文将QPP任务及其评估泛化为三种设置:(i) 单排序器多查询(SRMQ-PP),对应标准用例;(ii) 多排序器单查询(MRSQ-PP),评估QPP模型为查询选择最优排序器的能力;(iii) 多排序器多查询(MRMQ-PP),联合考虑所有查询-排序器组合的预测。结果表明:(a) QPP模型在不同任务间表现差异显著(SRMQ-PP vs. MRSQ-PP);(b) 预测查询的最佳排序器远比预测给定排序器下查询相对难度更困难。

原文摘要 · Abstract (English)

The traditional use-case of query performance prediction (QPP) is to identify which queries perform well and which perform poorly for a given ranking model. A more fine-grained and arguably more challenging extension of this task is to determine which ranking models are most effective for a given query. In this work, we generalize the QPP task and its evaluation into three settings: (i) SingleRanker MultiQuery (SRMQ-PP), corresponding to the standard use case; (ii) MultiRanker SingleQuery (MRSQ-PP), which evaluates a QPP model's ability to select the most effective ranker for a query; and (iii) MultiRanker MultiQuery (MRMQ-PP), which considers predictions jointly across all query ranker pairs. Our results show that (a) the relative effectiveness of QPP models varies substantially across tasks (SRMQ-PP vs. MRSQ-PP), and (b) predicting the best ranker for a query is considerably more difficult than predicting the relative difficulty of queries for a given ranker.

信息检索性能预测排序模型评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。