提出新指标T-Retrievability,更精准衡量检索系统对文档的公平暴露。
T-Retrievability: A Topic-Focused Approach to Measure Fair Document Exposure in Information Retrieval
- 按主题分组计算文档可检索性,再聚合得整体指标。
- 发现现有方法易受主题相关性干扰,导致公平性误判。
- 适合评估神经排序模型的文档曝光公平性。
可检索性是基于集合的统计量,用于衡量文档在特定排名截断范围内被检索的期望(倒数)排名。若一个集合中所有文档的可检索性得分分布均匀,则表明文档曝光公平。尽管已有研究使用可检索性分数来量化集合的曝光公平性,但本文利用可检索性分数的分布来衡量检索模型的曝光偏差。我们假设:整个集合中可检索性分数分布不均,可能并非反映曝光偏差,而是主题相关性的体现。为此,我们提出一种聚焦主题的局部可检索性度量方法,称为T-Retrievability(主题可检索性),即先在多个主题相关的文档组上计算可检索性分数,再将这些局部值聚合得到集合级别的统计量。通过该方法的分析,揭示了各类神经排序模型曝光特征的新见解。结果表明,这种局部化度量能更细致地理解曝光公平性,为评估信息检索系统中的文档可及性提供了更可靠的途径。
原文摘要 · Abstract (English)
Retrievability of a document is a collection-based statistic that measures its expected (reciprocal) rank of being retrieved within a specific rank cut-off. A collection with uniformly distributed retrievability scores across documents is an indicator of fair document exposure. While retrievability scores have been used to quantify the fairness of exposure for a collection, in our work, we use the distribution of retrievability scores to measure the exposure bias of retrieval models. We hypothesise that an uneven distribution of retrievability scores across the entire collection may not accurately reflect exposure bias but rather indicate variations in topical relevance. As a solution, we propose a topic-focused localised retrievability measure, which we call \textit{T-Retrievability} (topic-retrievability), which first computes retrievability scores over multiple groups of topically-related documents, and then aggregates these localised values to obtain the collection-level statistics. Our analysis using this proposed T-Retrievability measure uncovers new insights into the exposure characteristics of various neural ranking models. The findings suggest that this localised measure provides a more nuanced understanding of exposure fairness, offering a more reliable approach for assessing document accessibility in IR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。