arXiv:2410.01748cs.LG2024-10被引 30

测试大模型解小学数学题的连贯推理能力,发现多数模型存在严重逻辑断层。

Not All LLM Reasoners Are Created Equal

  • 用前后关联的数学题对测试模型连贯推理能力
  • 小模型和专用模型在组合题上性能下降超30%以上
  • 适合评估模型真实推理能力,尤其关注链式思维

我们研究了大语言模型(LLMs)在小学数学(GSM)问题求解中的深度推理能力。为此,我们评估模型在成对现有数学应用题上的表现,其中第二个问题的答案依赖于正确解答第一个问题。研究发现,大多数模型存在显著的推理差距:在组合题上的表现明显低于独立解答每道题的总和。这一差距在较小、更高效且专精数学的模型中尤为突出。此外,指令微调策略和代码生成对不同规模模型的影响各异,而仅在GSM数据集上微调可能导致任务过拟合。分析表明,推理差距并非源于测试集泄露,而是由于额外上下文干扰和第二步推理能力差所致。总体而言,尽管标准基准测试表现相似,但不同模型在推理能力上存在系统性差异。

原文摘要 · Abstract (English)

We study the depth of grade-school math (GSM) problem-solving capabilities of LLMs. To this end, we evaluate their performance on pairs of existing math word problems together so that the answer to the second problem depends on correctly answering the first problem. Our findings reveal a significant reasoning gap in most LLMs, that is performance difference between solving the compositional pairs and solving each question independently. This gap is more pronounced in smaller, more cost-efficient, and math-specialized models. Moreover, instruction-tuning recipes and code generation have varying effects across LLM sizes, while finetuning on GSM can lead to task overfitting. Our analysis indicates that large reasoning gaps are not because of test-set leakage, but due to distraction from additional context and poor second-hop reasoning. Overall, LLMs exhibit systematic differences in their reasoning abilities, despite what their performance on standard benchmarks indicates.

大模型推理数学能力链式思维模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。