用参考模型概率空间筛选高质量数据,少一半数据也能更好对齐人类偏好。
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
- 利用参考模型概率差异自动识别优质训练样本
- 仅用30%-50%数据,评测得分提升0.1至0.4
- 技术任务上效果提升达0.4至0.98,适合资源受限场景
直接偏好优化(DPO)已成为对齐语言模型与人类偏好的主流方法。近期研究发现,其效果高度依赖训练数据质量——优选与劣选回复间清晰的质量差异能显著提升学习性能。现有方法识别高质量样本需额外资源或外部模型。我们发现,参考模型的概率空间天然可检测高质量训练样本。基于此,提出一种采样策略,在仅使用不到一半(30%-50%)训练数据的前提下,于MT-Bench上实现一致提升(+0.1至+0.4)。在多个模型和超参数设置下,技术类任务(编程、数学、推理)表现大幅提升(+0.4至+0.98)。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has emerged as a de-facto approach for aligning language models with human preferences. Recent work has shown DPO's effectiveness relies on training data quality. In particular, clear quality differences between preferred and rejected responses enhance learning performance. Current methods for identifying and obtaining such high-quality samples demand additional resources or external models. We discover that reference model probability space naturally detects high-quality training samples. Using this insight, we present a sampling strategy that achieves consistent improvements (+0.1 to +0.4) on MT-Bench while using less than half (30-50%) of the training data. We observe substantial improvements (+0.4 to +0.98) for technical tasks (coding, math, and reasoning) across multiple models and hyperparameter settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。