解析偏好数据中的性能提升来源,发现模型差异和样本质量是关键。
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?

- 区分生成者级与样本级差异,分析偏好对齐中的有效信号。
- 增大生成模型能力差距可显著提升跨域推理表现。
- 按样本质量筛选数据,能实现更高效训练,适合模型优化者。
偏好优化方法如DPO和KTO广泛用于语言模型对齐,但其对下游推理能力的提升机制尚不明确。本文探究偏好对中影响推理模型性能的两类质量差异:生成者级差异(来自选择与拒绝生成路径的模型能力不同)和样本级差异(来自单个偏好对内生成内容的质量差异)。通过改变生成模型规模与类型研究生成者级差异,并利用大模型作为裁判在多个推理质量维度上评估生成轨迹,发现增大生成者级差异可稳定提升跨领域推理表现;而基于样本级差异筛选数据可实现更高效训练。结果表明,提升推理性能需双管齐下:构建偏好对时最大化生成者级差异,训练时利用样本级差异挑选高信息量样本。
原文摘要 · Abstract (English)
Preference optimization methods such as DPO and KTO are widely used for aligning language models, yet little is understood about what properties of preference data drive downstream reasoning gains. We ask: what aspects of a preference pair improve a reasoning model's performance on general reasoning tasks? We investigate two distinct notions of quality delta in preference data: generator-level delta, arising from the differences in capability between models that generate chosen and rejected reasoning traces, and sample-level delta, arising from differences in judged quality differences within an individual preference pair. To study generator-level delta, we vary the generator's scale and model family, and to study sample-level delta, we employ an LLM-as-a-judge to rate the quality of generated traces along multiple reasoning-quality dimensions. We find that increasing generator-level delta steadily improves performance on out-of-domain reasoning tasks and filtering data by sample-level delta can enable more data-efficient training. Our results suggest a twofold recipe for improving reasoning performance through preference optimization: maximize generator-level delta when constructing preference pairs and exploit sample-level delta to select the most informative training examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。