arXiv:2410.20147cs.LG2024-10被引 9

用生成流网络让大模型学会多解数学题,提升教育应用价值。

GFlowNet Fine-tuning for Diverse Correct Solutions in Mathematical Reasoning Tasks

  • 采用生成流网络(GFlowNet)训练模型,学习多样解题路径。
  • 在数学推理任务中,模型生成正确答案且中间步骤差异显著。
  • 适合需要多角度解题的教育场景或个性化辅导系统。

数学推理问题极具挑战性,通常需理解基本规律才能求解。尽管规律普遍,但解题路径因方法不同而异。在训练大语言模型(LLMs)时,掌握生成多种正确解法的能力对加速其在数学教育中的应用至关重要。为此,我们采用生成流网络(GFlowNet)对LLMs进行微调。与奖励最大化的强化学习(RL)不同,GFlowNet微调旨在通过训练使模型分布与奖励函数成比例,从而发现多样解法。在数值实验中,我们从准确率和多样性两个维度评估了GFlowNet微调与奖励最大化RL的效果。结果表明,GFlowNet微调能从多样化的中间推理步骤中导出正确最终答案,显著提升了生成替代解法的能力。

原文摘要 · Abstract (English)

Mathematical reasoning problems are among the most challenging, as they typically require an understanding of fundamental laws to solve. The laws are universal, but the derivation of the final answer changes depending on how a problem is approached. When training large language models (LLMs), learning the capability of generating such multiple solutions is essential to accelerate their use in mathematical education. To this end, we train LLMs using generative flow network (GFlowNet). Different from reward-maximizing reinforcement learning (RL), GFlowNet fine-tuning seeks to find diverse solutions by training the LLM whose distribution is proportional to a reward function. In numerical experiments, we evaluate GFlowNet fine-tuning and reward-maximizing RL in terms of accuracy and diversity. The results show that GFlowNet fine-tuning derives correct final answers from diverse intermediate reasoning steps, indicating the improvement of the capability of alternative solution generation.

数学推理生成流网络多解生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。