通过多样化数据提升大模型数学推理能力
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
- 设计结构化方法DTS,分解问题生成多样解题路径
- 在GSM8K和MATH上分别提升7.1%和4.2%准确率
- 计算开销仅增加3%,适合高效对齐训练
尽管偏好学习在对齐人类反馈方面取得进展,数学推理仍是大语言模型的持续挑战。本文研究偏好优化中数据多样化策略对数学推理能力的影响。评估了温度采样、思维链提示和蒙特卡洛树搜索三种常用数据生成方法,并提出新型结构化方法Diversified-ThinkSolve(DTS),系统性地将问题分解为多样化的推理路径。结果表明,通过有策略的数据多样化,模型可显著提升数学推理表现,在GSM8K上比基线提升7.1%,在MATH上提升4.2%。尽管性能优异,DTS的计算开销仅为基线的1.03倍,而MCTS成本接近五倍且收益更低。这表明,对多样化解题方式的结构化探索,比传统方法更有效于数学对齐。
原文摘要 · Abstract (English)
While recent advances in preference learning have enhanced alignment in human feedback, mathematical reasoning remains a persistent challenge. We investigate how data diversification strategies in preference optimization can improve the mathematical reasoning abilities of large language models (LLMs). We evaluate three common data generation methods: temperature sampling, Chain-of-Thought prompting, and Monte Carlo Tree Search (MCTS), and introduce Diversified-ThinkSolve (DTS), a novel structured approach that systematically decomposes problems into diverse reasoning paths. Our results show that with strategically diversified preference data, models can substantially improve mathematical reasoning performance, with the best approach yielding gains of 7.1% on GSM8K and 4.2% on MATH over the base model. Despite its strong performance, DTS incurs only a marginal computational overhead (1.03x) compared to the baseline, while MCTS is nearly five times more costly with lower returns. These findings demonstrate that structured exploration of diverse problem-solving methods creates more effective preference data for mathematical alignment than traditional approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。