arXiv:2410.21438cs.CLcs.LG2024-10被引 8

统一微调框架让大模型训练更高效,避免性能下降。

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

  • 用隐式奖励函数合并指令微调与对齐训练
  • 在指令遵循和事实性任务上显著提升性能
  • 适合需要高效后训练的大模型研究者

通过万亿级文本预训练,大语言模型获得文本生成能力。但为提升实用性并减少潜在危害,需依次进行指令微调(SFT)与对齐训练。由于两者目标和过程不同,某些任务性能可能下降。为此,我们提出统一微调(UFT),通过隐式奖励函数将SFT与对齐整合到单一训练阶段,使用相同目标和损失函数。实验表明,仅使用指令微调数据时,UFT优于传统SFT;当结合指令微调与对齐数据时,UFT有效防止任务性能退化,并在 extbf{ifeval}(指令遵循)和 extbf{truthful}(事实性)任务上表现更优,验证了UFT作为大模型后训练的有效高效范式。

原文摘要 · Abstract (English)

By pretraining on trillions of tokens, an LLM gains the capability of text generation. However, to enhance its utility and reduce potential harm, SFT and alignment are applied sequentially to the pretrained model. Because SFT and alignment have different objectives and underlying processes, performance on certain tasks can decline. To address this, we seamlessly introduce Unified Fine-Tuning (UFT), which integrates SFT and alignment into a single training stage using the same objective and loss functions through an implicit reward function. Our experimental results demonstrate that UFT outperforms SFT on instruction-tuning data alone. Moreover, when combining instruction-tuning data with alignment data, UFT effectively prevents the degradation on some tasks across these two stages and shows a clear advantage over sequentially applying SFT and alignment. This is evident in the significant improvements observed in the \textbf{ifeval} task for instruction-following and the \textbf{truthful} task for factuality. The proposed general fine-tuning framework UFT establishes an effective and efficient paradigm for LLM post-training.

大模型微调强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。