arXiv:2412.11699cs.CL2024-12ACL被引 2

通过多样化代码风格提升数学大模型推理能力

CoinMath: Harnessing the Power of Coding Instruction for Math LLMs

  • 用简洁注释、描述性命名和硬编码解法增强代码型推理
  • 新方法在数学基准上超越SOTA模型MAmmoTH
  • 适合想提升数学推理的LLM研究者和开发者

大型语言模型在解决数学问题上表现强劲,其中基于代码的解题方案尤为有效。然而,如何利用编程指令数据提升数学推理能力仍不明确。本研究探讨三个关键问题:(1) 不同数学代码推理风格对LLM学习效果的影响;(2) 通用领域编程指令能否提升性能;(3) 训练中融合文本与代码推理是否增强数学推理能力。结果表明,包含简洁注释、描述性命名和硬编码解法的代码型推理有益,而通用编程指令和文本推理的提升有限。基于此,提出CoinMath策略,通过多样化代码风格生成多种代码型推理。实验显示,CoinMath显著优于其基线模型MAmmoTH,该模型为当前最优数学大模型之一。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong performance in solving mathematical problems, with code-based solutions proving particularly effective. However, the best practice to leverage coding instruction data to enhance mathematical reasoning remains underexplored. This study investigates three key questions: (1) How do different coding styles of mathematical code-based rationales impact LLMs' learning performance? (2) Can general-domain coding instructions improve performance? (3) How does integrating textual rationales with code-based ones during training enhance mathematical reasoning abilities? Our findings reveal that code-based rationales with concise comments, descriptive naming, and hardcoded solutions are beneficial, while improvements from general-domain coding instructions and textual rationales are relatively minor. Based on these insights, we propose CoinMath, a learning strategy designed to enhance mathematical reasoning by diversifying the coding styles of code-based rationales. CoinMath generates a variety of code-based rationales incorporating concise comments, descriptive naming conventions, and hardcoded solutions. Experimental results demonstrate that CoinMath significantly outperforms its baseline model, MAmmoTH, one of the SOTA math LLMs.

数学推理代码生成LLM优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。