提出模型偏好对齐的统计度量方法,揭示模型排序与参考偏好的一致性。
Pairwise Reference Alignment as a Model-Level Ordinal Observable

- 定义模型得分函数与参考偏好对的一致性为序数可观测量
- 实证显示对齐度随模型规模和指令微调提升,且在不同数据子集间变化
- 提供可计算的估计器和统计边界,适用于奖励建模与对齐评估
成对偏好数据广泛用于语言模型评估与对齐,常用于模型排序、奖励建模或偏好优化。本文探讨一个更基础的测量问题:给定成对偏好的参考分布,当检验模型是否将优选回应排在拒斥回应之上时,实际估计的是何种模型级量?我们定义了由模型评分函数导出的成对参考对齐(pairwise reference alignment)作为序数可观测量。给定参考对分布 $P_{\mathrm{pair}}$ 在三元组 $(x,y^+,y^-)$ 上,以及标量模型得分 $S_M(x,y)$,对齐可观测量定义为模型诱导排序与参考偏好排序一致的概率。进一步定义中心化类似序参数的统计量,并讨论基于间隔的扩展。所得量可在独立采样假设下获得简单有限样本估计器和浓度界。本文不引入新基准,而是提供成对参考对齐的概念与统计框架,阐明参考对分布的作用,并区分一般序数可观测量与具体评分选择(如归一化对数似然或能量评分)。我们在 Qwen2.5 模型和 RewardBench 数据集上进行初步实证研究,结果表明所提统计量随模型规模和指令微调增加,并在不同参考对子集间呈现预测中的差异。
原文摘要 · Abstract (English)
Pairwise preference data is widely used in language-model evaluation and alignment, often for model ranking, reward modeling, or preference optimization. This note formulates a more basic measurement question: given a reference distribution of pairwise preferences, what model-level quantity is estimated when we test whether a model ranks preferred responses above rejected responses? We define pairwise reference alignment as an ordinal observable induced by a model scoring function. Given a reference pair distribution $P_{\mathrm{pair}}$ over triples $(x,y^+,y^-)$, and a scalar model score $S_M(x,y)$, we define the alignment observable as the probability that the model-induced ordering agrees with the reference preference ordering. We further define a centered order-parameter-like statistic and discuss a margin-based extension. The resulting quantities admit simple finite-sample estimators and concentration bounds under independent sampling assumptions. This note does not introduce a new benchmark. It provides a conceptual and statistical formulation for pairwise reference alignment, clarifies the role of the reference pair distribution, and distinguishes the general ordinal observable from scoring choices such as normalized log-probability or energy-based scores. We also provide an initial empirical study on Qwen2.5 models and RewardBench, where the proposed statistics increase with model size and instruction tuning and vary across reference-pair subsets as predicted by the formulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。