通过增强数据集提升多回复偏好优化效果,让大模型同时学习多个回答。
Multi-Response Preference Optimization with Augmented Ranking Dataset
- 用增强数据集改进偏好优化训练数据质量
- 支持同时学习多个回复,提升模型泛化能力
- 适合需要高质量对话生成的场景
大型语言模型(LLMs)近年来进展显著,新模型持续超越旧版本。这一进步得益于对多种训练机制的深入研究。其中,偏好优化通过融入人类偏好,在提升模型性能方面发挥了重要作用。然而,构建偏好优化数据集难度大,且优化过程对数据质量极为敏感。本文提出一种新颖的数据集增强方法,并引入基于多回复的偏好优化训练策略,实现对多个回复的并行学习,从而提高模型在复杂任务中的表现。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have been remarkable, with new models consistently surpassing their predecessors. These advancements are underpinned by extensive research on various training mechanisms. Among these, Preference Optimization has played a significant role in improving the performance of LLMs by incorporating human preferences into the training process. However, constructing preference optimization datasets is challenging and the optimization process is highly sensitive to the dataset quality. In this study, we propose a novel approach to augment Preference Optimization datasets. Additionally, we introduce a Multi-response-based Preference Optimization training method that enables the simultaneous learning of multiple responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。