测试大模型解组合数学题能力,发现GPT-4表现优于人类选手。
Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments
- 构建125道变体组合题,通过语义扰动检验模型泛化力。
- GPT-4在数学变体题上准确率超人类参赛者,达87.6%。
- 模型性能受题目表述影响大,人类却几乎不受干扰。
本文评估近期大型语言模型(LLMs)在组合数学问题上的求解能力。对比了LLaMA-2、LLaMA-3.1、GPT-4和Mixtral模型,以及具有奥数经验的中小学生与本科生。为此构建了Combi-Puzzles数据集,包含基于25个组合推理题的125道变体题,每题以五种形式呈现:通过对抗性添加、数值参数变化和语言模糊化系统性地改写题干。这些变体保持数学核心不变,旨在衡量模型的泛化能力,并提高问题在训练中未出现的可能性。实验发现,基于GPT-4的模型在生成正确答案方面优于其他所有模型,在数学变体题上表现显著优于人类,准确率达87.6%。同时,题干修改显著影响模型表现,而人类表现基本不受影响。
原文摘要 · Abstract (English)
In this paper we look at the ability of recent large language models (LLMs) at solving mathematical problems in combinatorics. We compare models LLaMA-2, LLaMA-3.1, GPT-4, and Mixtral against each other and against human pupils and undergraduates with prior experience in mathematical olympiads. To facilitate these comparisons we introduce the Combi-Puzzles dataset, which contains 125 problem variants based on 25 combinatorial reasoning problems. Each problem is presented in one of five distinct forms, created by systematically manipulating the problem statements through adversarial additions, numeric parameter changes, and linguistic obfuscation. Our variations preserve the mathematical core and are designed to measure the generalisability of LLM problem-solving abilities, while also increasing confidence that problems are submitted to LLMs in forms that have not been seen as training instances. We found that a model based on GPT-4 outperformed all other models in producing correct responses, and performed significantly better in the mathematical variation of the problems than humans. We also found that modifications to problem statements significantly impact the LLM's performance, while human performance remains unaffected.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。