arXiv:2412.20996cs.CL2024-12被引 2

让大模型重点学难题,提升数学推理能力

Plug-and-Play Training Framework for Preference Optimization

  • 根据输出分布动态加权样本,识别难例优先训练
  • 在数学推理任务上显著提升准确率,效果稳定
  • 可无缝接入多种优化方法,适合高精度场景

近期的偏好优化方法(如DPO)显著提升了大语言模型在对话和问答等任务中的表现。然而,现有方法未考虑训练样本难度差异,导致在高精度要求任务中表现平庸,尤其在数学推理方面。为此,我们提出一种新型训练框架,通过多轮采样分析输出分布,为不同样本分配权重,并将权重融入偏好优化过程。该即插即用方法使大模型在训练时更关注困难样本,提升学习效率。实验表明,该框架能与多种偏好优化方法无缝集成,在数学推理任务中实现持续性能提升。

原文摘要 · Abstract (English)

Recently, preference optimization methods such as DPO have significantly enhanced large language models (LLMs) in wide tasks including dialogue and question-answering. However, current methods fail to account for the varying difficulty levels of training samples during preference optimization, leading to mediocre performance in tasks with high accuracy requirements, particularly in mathematical reasoning. To address this limitation, we propose a novel training framework, which employs multiple sampling to analyze output distributions, assign different weights to samples, and incorporate these weights into the preference optimization process. This plug-and-play approach enables LLMs to prioritize challenging examples during training, improving learning efficiency. Experimental results demonstrate that our framework integrates seamlessly with various preference optimization methods and achieves consistent improvements in mathematical reasoning tasks.

偏好优化数学推理训练框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。