arXiv:2410.06961cs.CLcs.AI2024-10ICLR被引 38

用自生成偏好数据让大模型自我优化,省去人工标注

Self-Boosting Large Language Models with Synthetic Preference Data

  • 用自动生成提示和改进响应实现模型自我对齐
  • 迭代4次后指令遵循能力提升超22.1%,多任务表现增3.2~5.0分
  • 适合想降低标注成本、持续优化模型的研究者

通过与人类偏好对齐,大语言模型在生成诚实、无害、有帮助的回复方面取得了显著进展。然而,获取高质量偏好数据是一项资源密集且需要创造力的过程,尤其在持续改进大模型时更为困难。我们提出SynPO,一种利用合成偏好数据进行模型对齐的自增强范式。SynPO采用迭代机制:自提示生成器创造多样化提示,响应改进器逐步优化模型输出。该方法使大模型能自主学习自身输出的生成奖励,无需大规模提示标注和人工偏好数据。经过四轮SynPO迭代,Llama3-8B和Mistral-7B在指令遵循能力上显著提升,在AlpacaEval 2.0和ArenaHard上分别实现超过22.1%的胜率提升。同时,模型在各类任务上的通用性能也得到改善,体现在知名Open LLM Leaderboard上平均得分提升3.2至5.0分。

原文摘要 · Abstract (English)

Through alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard.

大模型对齐自生成数据指令优化自动化训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。