提出新评估框架,让查询性能预测更贴近实际搜索应用。
Beyond Correlations: A Downstream Evaluation Framework for Query Performance Prediction
- 用多个排序器的文档列表中预测质量分布作为融合先验
- 加权融合比无权重提升超4.5%效果
- 发现传统相关性评估无法反映真实应用价值
当前查询性能预测(QPP)评估普遍依赖估计与真实检索质量之间的整体相关性,但该方法既无法衡量单个查询的预测效果,也未对接下游任务,导致高相关性结果未必适用于实际检索决策。本文提出一种面向下游应用的评估框架:利用多个排序器生成的前Top文档中预测质量的分布作为加权融合的先验信息。一方面,该分布若贴近真实质量分布,说明预测准确;另一方面,其在融合中的实际表现则体现预测器在检索系统中做出有效决策的能力。实验表明,使用QPP估计进行加权融合可显著提升效果,较无权重的CombSUM和RRF策略提升超过4.5%;同时揭示:传统相关性指标与下游实际效果关联性较差。
原文摘要 · Abstract (English)
The standard practice of query performance prediction (QPP) evaluation is to measure a set-level correlation between the estimated retrieval qualities and the true ones. However, neither this correlation-based evaluation measure quantifies QPP effectiveness at the level of individual queries, nor does this connect to a downstream application, meaning that QPP methods yielding high correlation values may not find a practical application in query-specific decisions in an IR pipeline. In this paper, we propose a downstream-focussed evaluation framework where a distribution of QPP estimates across a list of top-documents retrieved with several rankers is used as priors for IR fusion. While on the one hand, a distribution of these estimates closely matching that of the true retrieval qualities indicates the quality of the predictor, their usage as priors on the other hand indicates a predictor's ability to make informed choices in an IR pipeline. Our experiments firstly establish the importance of QPP estimates in weighted IR fusion, yielding substantial improvements of over 4.5% over unweighted CombSUM and RRF fusion strategies, and secondly, reveal new insights that the downstream effectiveness of QPP does not correlate well with the standard correlation-based QPP evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。