arXiv:2602.10286cs.LG2026-02被引 2

揭示偏好学习在真实数据中实际恢复的是什么

What Does Preference Learning Recover from Pairwise Comparison Data?

  • 从成对比较数据出发,定义条件偏好分布来刻画偏好信息
  • 发现BT模型适用的精确条件及影响样本效率的关键因素
  • 为理解偏好学习本质提供数据驱动的新视角

成对偏好学习是机器学习的核心,广泛应用于语言模型与人类偏好对齐。典型数据集由三元组 (x, y+, y−) 构成,表示在上下文 x 下响应 y+ 被偏好于 y−。主流方法采用Bradley-Terry(BT)模型,将偏好概率建模为潜在评分差值的函数。标准做法假设数据服从该模型并据此学习潜在评分。然而真实数据可能违背此假设,且当前尚不清楚在非理想情况下BT学习究竟恢复了什么。本文从三元组比较数据出发,通过条件偏好分布(CPRD)形式化地刻画其编码的偏好信息。我们给出了BT模型适用于建模CPRD的精确条件,并识别出影响样本效率的关键因素——边界(margin)和连通性(connectivity)。这些结果共同为理解偏好学习实际恢复的内容提供了数据驱动的基础。

原文摘要 · Abstract (English)

Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets $(x, y^+, y^-)$, where response $y^+$ is preferred over response $y^-$ for context $x$. The Bradley--Terry (BT) model is the predominant approach, modeling preference probabilities as a function of latent score differences. Standard practice assumes data follows this model and learns the latent scores accordingly. However, real data may violate this assumption, and it remains unclear what BT learning recovers in such cases. Starting from triplet comparison data, we formalize the preference information it encodes through the conditional preference distribution (CPRD). We give precise conditions for when BT is appropriate for modeling the CPRD, and identify factors governing sample efficiency -- namely, margin and connectivity. Together, these results offer a data-centric foundation for understanding what preference learning actually recovers.

偏好学习机器学习统计建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。