用代码生成偏好数据,提升大模型推理能力
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning
- 用公开代码生成大量代码-偏好对,预训练偏好模型
- 在GSM8K、MATH等任务上显著提升模型推理性能
- 适合研究高效奖励建模与大模型推理增强的学者
大型语言模型在自然语言理解与生成方面取得显著进展,主要得益于可扩展的预训练和先进的微调技术。然而,通过人类反馈强化学习(RLHF)提升模型推理能力仍面临挑战,主要源于高质量偏好数据稀缺——这类数据需人工标注且耗时费力,对奖励模型微调至关重要。为缓解此问题,我们提出CodePMP,一种可扩展的偏好模型预训练(PMP)流程,利用公开高质量源代码构建大规模合成代码-偏好对。CodePMP通过在海量合成代码-偏好对上预训练偏好模型,提升了奖励模型微调效率。我们在数学推理任务(GSM8K、MATH)和逻辑推理任务(ReClor、LogiQA2.0)上评估,结果一致显示模型推理能力显著提升,凸显了可扩展偏好模型预训练在高效奖励建模中的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) have made significant progress in natural language understanding and generation, driven by scalable pretraining and advanced finetuning. However, enhancing reasoning abilities in LLMs, particularly via reinforcement learning from human feedback (RLHF), remains challenging due to the scarcity of high-quality preference data, which is labor-intensive to annotate and crucial for reward model (RM) finetuning. To alleviate this issue, we introduce CodePMP, a scalable preference model pretraining (PMP) pipeline that utilizes a large corpus of synthesized code-preference pairs from publicly available high-quality source code. CodePMP improves RM finetuning efficiency by pretraining preference models on large-scale synthesized code-preference pairs. We evaluate CodePMP on mathematical reasoning tasks (GSM8K, MATH) and logical reasoning tasks (ReClor, LogiQA2.0), consistently showing significant improvements in reasoning performance of LLMs and highlighting the importance of scalable preference model pretraining for efficient reward modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。