arXiv:2505.10527cs.CL2025-05被引 7

发现偏好模型性能随规模增长的幂律规律,构建可扩展的人类偏好统一模型。

WorldPM: Scaling Human Preference Modeling

  • 基于大规模公开论坛数据,训练1.5B至72B参数模型,探索偏好建模的规模效应。
  • 在对抗性与客观性任务上,模型性能随数据和参数量增加显著提升,最大增益超5%。
  • 适合用于强化学习人类反馈(RLHF)系统,尤其对小样本偏好数据有强泛化能力。

受语言建模中测试损失随模型与数据规模呈幂律变化的启发,我们发现偏好建模也存在类似规律。本文提出世界偏好建模(WorldPM),强调其规模化潜力,其中世界偏好代表人类偏好的统一表征。我们在涵盖多样用户群体的公开论坛中收集偏好数据,并在1500万规模的数据上,对1.5B至72B参数的模型进行广泛训练。观察到不同评估指标下的显著差异:(1) 对抗性指标(识别欺骗性特征能力)随训练数据和基础模型规模增加而持续提升;(2) 客观性指标(具有明确答案的知识任务)在大语言模型中表现出涌现行为,凸显WorldPM的可扩展性;(3) 主观性指标(少数人类或AI的主观偏好)未呈现明显增长趋势。通过7个基准、20个子任务的评估,验证了WorldPM作为偏好微调基础的有效性,其在7K、100K及800K样本的不同规模人类偏好数据集上均实现广泛性能提升,多个关键子任务提升超过5%。将其集成至内部RLHF流程后,在自研与公共评估集上均有显著改进,自研评估中提升达4%至8%。

原文摘要 · Abstract (English)

Motivated by scaling laws in language modeling that demonstrate how test loss scales as a power law with model and dataset sizes, we find that similar laws exist in preference modeling. We propose World Preference Modeling$ (WorldPM) to emphasize this scaling potential, where World Preference embodies a unified representation of human preferences. In this paper, we collect preference data from public forums covering diverse user communities, and conduct extensive training using 15M-scale data across models ranging from 1.5B to 72B parameters. We observe distinct patterns across different evaluation metrics: (1) Adversarial metrics (ability to identify deceptive features) consistently scale up with increased training data and base model size; (2) Objective metrics (objective knowledge with well-defined answers) show emergent behavior in larger language models, highlighting WorldPM's scalability potential; (3) Subjective metrics (subjective preferences from a limited number of humans or AI) do not demonstrate scaling trends. Further experiments validate the effectiveness of WorldPM as a foundation for preference fine-tuning. Through evaluations on 7 benchmarks with 20 subtasks, we find that WorldPM broadly improves the generalization performance across human preference datasets of varying sizes (7K, 100K and 800K samples), with performance gains exceeding 5% on many key subtasks. Integrating WorldPM into our internal RLHF pipeline, we observe significant improvements on both in-house and public evaluation sets, with notable gains of 4% to 8% in our in-house evaluations.

偏好建模规模化RLHF大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。