二元反馈比序数比较更准,能更快恢复排序。
When Less Is More: Binary Feedback Can Outperform Ordinal Comparisons in Ranking Recovery
- 用广义可加模型统一建模序数与二元对比数据
- 理论证明二元反馈的排序误差收敛速度更快
- 适合偏好学习、推荐系统等需要高效排序的场景
成对比较数据在排序与偏好学习中至关重要。尽管序数比较看似比二元比较信息更丰富,本文挑战这一传统认知。我们提出一个通用参数化框架,用于建模无平局的序数成对比较。该模型采用广义可加结构,包含链接函数(量化两项目间偏好差异)和模式函数(控制序数响应分布)。此框架将经典二元比较模型作为特例,将二元响应视为序数数据的二值化版本。理论上证明:在计数算法下,二元比较的排序误差具有更快的指数收敛速率,优于序数数据。进一步通过由模式函数决定的信噪比(SNR)刻画二元与序数数据间的显著性能差距,识别出最小化SNR的模式函数,最大化二值化的收益。大量模拟实验与真实数据集MovieLens上的应用验证了理论结果。
原文摘要 · Abstract (English)
Paired comparison data, where users evaluate items in pairs, play a central role in ranking and preference learning tasks. While ordinal comparison data intuitively offer richer information than binary comparisons, this paper challenges that conventional wisdom. We propose a general parametric framework for modeling ordinal paired comparisons without ties. The model adopts a generalized additive structure, featuring a link function that quantifies the preference difference between two items and a pattern function that governs the distribution over ordinal response levels. This framework encompasses classical binary comparison models as special cases, by treating binary responses as binarized versions of ordinal data. Within this framework, we show that binarizing ordinal data can significantly improve the accuracy of ranking recovery. Specifically, we prove that under the counting algorithm, the ranking error associated with binary comparisons exhibits a faster exponential convergence rate than that of ordinal data. Furthermore, we characterize a substantial performance gap between binary and ordinal data in terms of a signal-to-noise ratio (SNR) determined by the pattern function. We identify the pattern function that minimizes the SNR and maximizes the benefit of binarization. Extensive simulations and a real application on the MovieLens dataset further corroborate our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。