arXiv:2602.08033cs.LG2026-02

混合对比与评分可更高效准确地评估对象排序。

The Benefits of Diversity: Combining Comparisons and Ratings for Efficient Scoring

  • 统一建模对比与评分两种偏好信息,提升评分精度。
  • 在模型不匹配情况下仍能恢复准确分数,且保证单调性与鲁棒性。
  • 适合需要精准排名头部实体的场景,如推荐系统、筛选任务。

人类应单独评价对象,还是进行对比评价?这一问题长期存在争议。本文发现,将两种方式结合使用,反而优于单一方式。我们提出SCoRa(基于对比与评分的打分)模型,一种统一的概率框架,可同时利用两种信号学习偏好。理论证明,SCoRa的极大后验估计器具有良好性质:具备单调性与鲁棒性保障。实验显示,即使在模型不匹配时,该方法仍能准确恢复真实分数。最值得注意的是,在一个现实场景中,结合两种信号的表现优于仅用任一形式,尤其在头部实体排序至关重要的情况下。由于多形式信号普遍存在,SCoRa为偏好学习提供了一个灵活通用的基础。

原文摘要 · Abstract (English)

Should humans be asked to evaluate entities individually or comparatively? This question has been the subject of long debates. In this work, we show that, interestingly, combining both forms of preference elicitation can outperform the focus on a single kind. More specifically, we introduce SCoRa (Scoring from Comparisons and Ratings), a unified probabilistic model that allows to learn from both signals. We prove that the MAP estimator of SCoRa is well-behaved. It verifies monotonicity and robustness guarantees. We then empirically show that SCoRa recovers accurate scores, even under model mismatch. Most interestingly, we identify a realistic setting where combining comparisons and ratings outperforms using either one alone, and when the accurate ordering of top entities is critical. Given the de facto availability of signals of multiple forms, SCoRa additionally offers a versatile foundation for preference learning.

偏好学习评分模型对比评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。