解决低精度检索中的评分歧义问题,提升评估可靠性。
Reliable Evaluation Protocol for Low-Precision Retrieval
- 高精度评分:最后一轮打分升至高精度,低成本化解并列
- 并列感知指标:量化并列带来的排序不确定性和偏差
- 适合关注低精度模型评估可靠性的研究者
降低模型参数与计算的数值精度被广泛用于提升检索系统效率。但在低精度下计算查询与文档的相关性得分时,因精度降低导致虚假并列现象,使结果对并列处理方式高度敏感,评估可靠性下降。为此,我们提出一种更鲁棒的检索评估协议:(1) 高精度评分(HPS),将最终打分步骤提升至高精度,以最小计算成本化解并列候选;(2) 并列感知检索指标(TRM),报告预期得分、得分范围与偏差,量化并列候选的排序不确定性。在两个检索数据集上,使用多种模型与三种打分函数的实验表明,HPS显著降低由并列引发的不稳定性,TRM能准确恢复预期指标值。该组合实现了更低精度检索下更一致、可靠的评估体系。
原文摘要 · Abstract (English)
Lowering the numerical precision of model parameters and computations is widely adopted to improve the efficiency of retrieval systems. However, when computing relevance scores between the query and documents in low-precision, we observe spurious ties due to the reduced granularity. This introduces high variability in the results based on tie resolution, making the evaluation less reliable. To address this, we propose a more robust retrieval evaluation protocol designed to reduce score variation. It consists of: (1) High-Precision Scoring (HPS), which upcasts the final scoring step to higher precision to resolve tied candidates with minimal computational cost; and (2) Tie-aware Retrieval Metrics (TRM), which report expected scores, range, and bias to quantify order uncertainty of tied candidates. Our experiments test multiple models with three scoring functions on two retrieval datasets to demonstrate that HPS dramatically reduces tie-induced instability, and TRM accurately recovers expected metric values. This combination enables a more consistent and reliable evaluation system for lower-precision retrievals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。