在无法计算召回率时,如何准确评估检索质量?
How important is Recall for Measuring Retrieval Quality?
- 用大模型判断生成结果质量,间接衡量检索效果
- 在仅2-15个相关文档的场景下验证指标有效性
- 提出无需知道总相关文档数的新评估方法
在大规模且动态更新的知识库中,查询的真实相关文档总数通常未知,导致召回率无法计算。本文通过测量检索质量指标与基于大模型的响应质量判断之间的相关性,评估多种应对策略的有效性。实验在多个数据集上进行,相关文档数量较少(2-15篇)。同时,提出一种简单有效的检索质量度量方法,无需预先知晓相关文档总数即可取得良好表现。
原文摘要 · Abstract (English)
In realistic retrieval settings with large and evolving knowledge bases, the total number of documents relevant to a query is typically unknown, and recall cannot be computed. In this paper, we evaluate several established strategies for handling this limitation by measuring the correlation between retrieval quality metrics and LLM-based judgments of response quality, where responses are generated from the retrieved documents. We conduct experiments across multiple datasets with a relatively low number of relevant documents (2-15). We also introduce a simple retrieval quality measure that performs well without requiring knowledge of the total number of relevant documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。