打造统一工具箱,让公平多样检索算法可比可测。
FairDiverse: A Comprehensive Toolkit for Fair and Diverse Information Retrieval Algorithms
- 整合预处理、内嵌、后处理三类方法,覆盖检索全流程。
- 支持16个基模型、28种算法在搜索与推荐任务上的评测。
- 开源可扩展,助力研究者快速开发并公平对比新模型。
现代信息检索(IR)不仅需要高精度,还需兼顾公平性与多样性以维持健康生态。尽管已有多种数据集、算法与评估框架,但因评测指标、数据集和实验设置不一,导致算法比较困难。为此,我们提出 FairDiverse——一个开源、标准化的 IR 工具箱。该工具箱支持在检索流程不同阶段集成预处理、内嵌、后处理等公平与多样性方法,涵盖16个基模型和28种算法,在搜索与推荐两类核心任务上建立全面基准。其高度可扩展性提供多接口支持,帮助研究者快速开发并公平评估新模型。项目已开源,地址:https://github.com/XuChen0427/FairDiverse。
原文摘要 · Abstract (English)
In modern information retrieval (IR). achieving more than just accuracy is essential to sustaining a healthy ecosystem, especially when addressing fairness and diversity considerations. To meet these needs, various datasets, algorithms, and evaluation frameworks have been introduced. However, these algorithms are often tested across diverse metrics, datasets, and experimental setups, leading to inconsistencies and difficulties in direct comparisons. This highlights the need for a comprehensive IR toolkit that enables standardized evaluation of fairness- and diversity-aware algorithms across different IR tasks. To address this challenge, we present FairDiverse, an open-source and standardized toolkit. FairDiverse offers a framework for integrating fair and diverse methods, including pre-processing, in-processing, and post-processing techniques, at different stages of the IR pipeline. The toolkit supports the evaluation of 28 fairness and diversity algorithms across 16 base models, covering two core IR tasks (search and recommendation) thereby establishing a comprehensive benchmark. Moreover, FairDiverse is highly extensible, providing multiple APIs that empower IR researchers to swiftly develop and evaluate their own fairness and diversity aware models, while ensuring fair comparisons with existing baselines. The project is open-sourced and available on https://github.com/XuChen0427/FairDiverse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。