arXiv:2410.16586cs.AIcs.IR2024-10被引 5

用少量偏好数据高效优化大模型,提升对齐效果

Optimizing LLMs with Direct Preferences: A Data Efficiency Perspective

  • 通过直接偏好优化,减少对大量标注数据依赖
  • 混合多样化数据集可显著提升模型表现
  • 对话类提示训练效果优于问答类提示

对齐大型语言模型(LLM)输出与人类偏好(如通过人类反馈强化学习,即RLHF)对于其在真实场景中的有效性至关重要。尽管在对齐技术方面取得显著进展,但不同类型的偏好数据对模型性能的影响尚未系统研究。本文探讨了直接偏好优化(DPO)在微调预训练大模型时的可扩展性、数据效率与有效性,旨在降低对大规模昂贵偏好数据的依赖。我们(1)系统比较了使用不同比例的综合偏好判断数据集时模型的表现,绘制出DPO的性能提升曲线,并评估其在数据受限环境下的效果;(2)为选择性使用偏好数据提供了优化策略。研究发现,增加训练数据量通常能提升并稳定模型性能;多种数据集组合使用可显著增强模型有效性;此外,使用对话类提示训练的模型表现优于问答类提示训练的模型。

原文摘要 · Abstract (English)

Aligning the output of Large Language Models (LLMs) with human preferences (e.g., by means of reinforcement learning with human feedback, or RLHF) is essential for ensuring their effectiveness in real-world scenarios. Despite significant advancements in LLM alignment techniques, the impact of different type of preference data on model performance has yet to be systematically explored. In this study, we investigate the scalability, data efficiency, and effectiveness of Direct Preference Optimization (DPO) in fine-tuning pre-trained LLMs, aiming to reduce their dependency on extensive amounts of preference data, which is expensive to collect. We (1) systematically compare the performance of models fine-tuned with varying percentages of a combined preference judgement dataset to define the improvement curve of DPO and assess its effectiveness in data-constrained environments; and (2) provide insights for the development of an optimal approach for selective preference data usage. Our study reveals that increasing the amount of data used for training generally enhances and stabilizes model performance. Moreover, the use of a combination of diverse datasets significantly improves model effectiveness. Furthermore, when models are trained separately using different types of prompts, models trained with conversational prompts outperformed those trained with question answering prompts.

大模型对齐偏好优化数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。