arXiv:2409.00131cs.CLcs.AI2024-09

用逻辑相似度检索题库,让小模型也能解数学应用题

Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems

  • 通过逻辑与语义双重相似度筛选参考题,构建高质量训练样本
  • 在SVAMP上比思维链提升15.8%,GSM8K上提升21.5%
  • 适用于资源有限但需高推理能力的轻量级模型场景

本研究聚焦于提升轻量级大语言模型在数学推理任务中的表现。提出一种测量数学逻辑相似性的新方法,并设计自动筛选机制,构建融合语义与逻辑相似性的参考题集。通过精心设计的正负例提示,引导模型采用正确的推理逻辑。据我们所知,这是首次将检索增强生成应用于数学问题求解。实验结果表明,该方法在SVAMP数据集上较思维链方法提升15.8%,在GSM8K数据集上提升21.5%。进一步将该方法应用于1750亿参数的大模型,性能达到当前最优水平。最后,对推理过程中的错误进行分析,为未来大模型推理研究提供宝贵洞见。

原文摘要 · Abstract (English)

This study focuses on improving the performance of lightweight Large Language Models (LLMs) in mathematical reasoning tasks. We introduce a novel method for measuring mathematical logic similarity and design an automatic screening mechanism to construct a set of reference problems that integrate both semantic and logical similarity. By employing carefully crafted positive and negative example prompts, we guide the model towards adopting sound reasoning logic. To the best of our knowledge, this is the first attempt to utilize retrieval-enhanced generation for mathematical problem-solving. Experimental results demonstrate that our method achieves a 15.8% improvement over the Chain of Thought approach on the SVAMP dataset and a 21.5 % improvement on the GSM8K dataset. Further application of this method to a large-scale model with 175 billion parameters yields performance comparable to the best results on both aforementioned datasets. Finally, we conduct an analysis of errors during the reasoning process, providing valuable insights and directions for future research on reasoning tasks using large language models.

数学推理轻量模型检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。