arXiv:2606.16970cs.IR2026-06

分析随机排序带来的效果波动风险,揭示其与初始召回点分布的关系。

A Theoretical Framework for Risk Analysis of Stochastic Rankers

  • 从初始检索列表的召回点分布出发,理论推导随机重排的风险上限。
  • 实验证明理论预测的DCG变化与实际观测值高度吻合。
  • 适合关注排序公平性与稳定性研究的学者参考。

与追求顶部相关性最大化的确定性排序不同,随机排序策略通过估计排列分布并从中采样,以实现多样或公平的曝光。这类策略通常在重排后评估期望有效性。然而,其固有的随机性引发了一个基础但未被充分探索的前置问题:在应用随机重排前,检索有效性可能产生的最坏情况变化有多大?本文提出了重排风险的理论分析,定义为从随机重排策略中采样一个排列作用于固定检索列表时,导致的折扣累积收益(DCG)最大绝对变化。我们推导出该风险由初始检索列表中的召回点分布决定。在TREC Fairness 2022赛道提交的采用随机重排策略的运行结果上进行实验,证实理论预测的有效性变化与实际观察到的DCG变化极为接近。

原文摘要 · Abstract (English)

Different from deterministic rankers that seek to maximize relevance at top ranks, stochastic ranking policies instead estimate distributions over permutations, from which rankings are sampled, towards obtaining diversified or fair exposure. Such policies are commonly evaluated in terms of expected effectiveness postreranking. However, the randomness inherent in these policies gives rise to a fundamental but under-explored ex ante question: prior to applying stochastic reranking, how large can the induced variation in retrieval effectiveness be in the worst case? This paper presents a theoretical analysis of reranking risk, defined as the maximum absolute change in discounted cumulative gain (DCG) resulting from a permutation sampled from a stochastic reranking policy applied to a fixed retrieved list.We derive that this risk is governed by the distribution of the recall points in the initial retrieved list. We conduct experiments on submitted runs from the TREC Fairness 2022 track that employ stochastic reranking policies and empirically demonstrate that the effectiveness variations predicted by our theory closely approximate the observed changes in DCG.

排序风险随机重排公平性理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。