提升开源模型在中英文数学推理能力,效果超越Gemini。
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
- 采用渐进式训练与分解策略,优化模型解题流程。
- WizardMath 7B 在英语上比Gemini高6%准确率,中文持平。
- 双语混合训练实现跨语言数学推理,适合资源受限场景。
大型语言模型在语言任务中表现优异,但在非英语语言如印地语的数学推理方面表现不佳。本研究旨在提升小型、资源高效的开源大模型在印地语和英语中的数学推理能力。我们评估了OpenHathi 7B、LLaMA-2 7B、WizardMath 7B、Mistral 7B、LLeMMa 7B、MAmmoTH 7B、Gemini Pro和GPT-4等模型,采用零样本、少样本链式思维(CoT)方法及监督微调。提出课程学习策略,逐步在更难问题上训练;引入新颖的分解策略简化复杂算术;设计结构化解题流程分阶段求解。实验显示显著性能提升:WizardMath 7B在英语数据集上准确率比Gemini高出+6%,在印地语数据集上达到与Gemini相当水平。双语混合训练方法可取得与单语模型相当的效果,证明模型具备同时学习双语数学推理的能力。研究展示了改进开源大模型数学推理潜力的可能性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in linguistic tasks but struggle with mathematical reasoning, particularly in non English languages like Hindi. This research aims to enhance the mathematical reasoning skills of smaller, resource efficient open-source LLMs in both Hindi and English. We evaluate models like OpenHathi 7B, LLaMA-2 7B, WizardMath 7B, Mistral 7B, LLeMMa 7B, MAmmoTH 7B, Gemini Pro, and GPT-4 using zero-shot, few-shot chain-of-thought (CoT) methods, and supervised fine-tuning. Our approach incorporates curriculum learning, progressively training models on increasingly difficult problems, a novel Decomposition Strategy to simplify complex arithmetic operations, and a Structured Solution Design that divides solutions into phases. Our experiments result in notable performance enhancements. WizardMath 7B exceeds Gemini's accuracy on English datasets by +6% and matches Gemini's performance on Hindi datasets. Adopting a bilingual approach that combines English and Hindi samples achieves results comparable to individual language models, demonstrating the capability to learn mathematical reasoning in both languages. This research highlights the potential for improving mathematical reasoning in open-source LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。