arXiv:2410.08068cs.CLcs.AI2024-10被引 3

模仿教师教学过程,提升大模型数学推理能力

Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models

  • 用教学式提示框架引入概念、定理和相似题解法
  • 在4个数学基准上达新最好成绩,最高提升7.2%
  • 适合需要增强逻辑推理的AI研究与教育应用

大语言模型在多个领域表现优异,但在算术推理任务上仍有不足。现有提示设计方法虽有效,但忽略了解决多数算术问题所需的关键概念、定理和技巧。为此,我们提出一种教学启发式集成框架,模拟教师指导学生的过程,向模型注入必要概念、相关定理及类似题目的解法思路,从而提升推理能力。此外,我们构建了两个带详细解析与答案的中文数据集:MathMC 和 MathToF。在九个基准上的实验表明,该方法显著提升模型推理准确率。使用 GPT-4 与本框架,在 AddSub、SVAMP、Math23K 和 AQuA 四个数学基准上分别达到 98.2%(+3.3%)、93.9%(+0.2%)、94.3%(+7.2%)和 81.1%(+1.2%)的准确率,创下新纪录。代码与数据已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit impressive performance across various domains but still struggle with arithmetic reasoning tasks. Recent work shows the effectiveness of prompt design methods in enhancing reasoning capabilities. However, these approaches overlook crucial requirements for prior knowledge of specific concepts, theorems, and tricks to tackle most arithmetic reasoning problems successfully. To address this issue, we propose a novel and effective Teaching-Inspired Integrated Framework, which emulates the instructional process of a teacher guiding students. This method equips LLMs with essential concepts, relevant theorems, and similar problems with analogous solution approaches, facilitating the enhancement of reasoning abilities. Additionally, we introduce two new Chinese datasets, MathMC and MathToF, both with detailed explanations and answers. Experiments are conducted on nine benchmarks which demonstrates that our approach improves the reasoning accuracy of LLMs. With GPT-4 and our framework, we achieve new state-of-the-art performance on four math benchmarks (AddSub, SVAMP, Math23K and AQuA) with accuracies of 98.2% (+3.3%), 93.9% (+0.2%), 94.3% (+7.2%) and 81.1% (+1.2%). Our data and code are available at https://github.com/SallyTan13/Teaching-Inspired-Prompting.

推理增强提示工程数学问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。