用伯莱多-泰瑞模型公平比较推荐系统,避免数据特性干扰排名。
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies

- 基于伯莱多-泰瑞模型构建数据驱动的算法排名方法。
- 排名结果受数据稀疏性、规模等关键统计量影响显著。
- 可预测未见数据集上的排名,适合算法选型与跨场景评估。
推荐算法的排名是一项挑战性任务,因模型性能对数据集特性(如稀疏性、序列结构、规模)敏感。直接聚合性能指标(如在多个基准上平均NDCG)可能导致误导性排名,影响实际选择。为此,我们提出一种基于伯莱多-泰瑞(Bradley-Terry, BT)模型的新颖数据驱动排名方法。我们证明所得排名依赖于关键数据集统计量。此外,提出一种新指标评估排名一致性,并验证了该方法对不完整数据的鲁棒性。最后,引入基于BT树和含协变量的BT模型的扩展框架,实现无需运行模型即可在未见数据集上进行算法排名的特定数据集方法。
原文摘要 · Abstract (English)
The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale. This drives a demand for a proper methodology for fair comparison between algorithms. Naive aggregation of performance metrics (e.g., averaging NDCG over benchmarks) can yield misleading rankings, undermining practical selection. To address this problem, we introduce a novel, data-driven ranking methodology based on Bradley-Terry (BT) model. We demonstrate that the obtained ranking depends on key dataset statistics. Additionally, we propose a novel metric for evaluating ranking consistency and demonstrate robustness of our ranking to incomplete data. Finally, we introduce a dataset-specific methodology for ranking algorithms on unseen datasets without running the models, relying on extensions of the Bradley-Terry framework, including BT trees and BT models with covariates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。