推理重排器对搜索公平性无显著影响,现有模型延续原始排序的公平性特征。
Does Reasoning Make Search More Fair? Comparing Fairness in Reasoning and Non-Reasoning Rerankers
- 对比六种重排模型,评估推理与非推理方法在公平性上的差异。
- 公平性指标AWRF稳定在0.33-0.35,不受相关性变化影响(nDCG 0.247-1.000)。
- 地理属性存在持续公平性差距,适合关注公平性研究的团队参考。
尽管推理重排器(如Rank1)在提升排序相关性方面表现优异,但其在其他检索质量如公平性方面的表现尚不明确。我们首次系统比较了推理与非推理重排器的公平性。基于TREC 2022公平排名赛道数据集,我们在多种检索设置和人口属性下评估了六种重排模型。结果表明,推理既未改善也未损害公平性,其公平性度量注意力加权排名公平性(AWRF)在所有模型中保持稳定(0.33–0.35),而相关性指标nDCG变化范围为0.247–1.000。按人口属性分解分析显示,无论模型架构如何,地理属性始终存在公平性差距。这说明未来将推理模型专门设计为感知公平性属性,可能带来改进,当前实现仍继承输入排序的公平性特征。
原文摘要 · Abstract (English)
While reasoning rerankers, such as Rank1, have demonstrated strong abilities in improving ranking relevance, it is unclear how they perform on other retrieval qualities such as fairness. We conduct the first systematic comparison of fairness between reasoning and non-reasoning rerankers. Using the TREC 2022 Fair Ranking Track dataset, we evaluate six reranking models across multiple retrieval settings and demographic attributes. Our findings demonstrate reasoning neither improve nor harm fairness compared to non-reasoning approaches. Our fairness metric, Attention-Weighted Rank Fairness (AWRF) remained stable (0.33-0.35) across all models, even as relevance varies substantially (nDCG 0.247-1.000). Demographic breakdown analysis revealed fairness gaps for geographic attributes regardless of model architecture. These results indicate that future work in specializing reasoning models to be aware of fairness attributes could lead to improvements, as current implementations preserve the fairness characteristics of their input ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。