构建复杂算术证明的评估框架,揭示大模型在难题上的推理衰减现象
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
- 用可配置的生成框架构造任意复杂算术证明问题
- 模型性能随证明树深度和广度增加显著下降,非线性结构更难应对
- 适合研究模型推理泛化能力,尤其关注复杂逻辑链的鲁棒性
大型语言模型(LLMs)在算术应用题上表现良好,但对其推广到更复杂问题的能力了解有限。这主要受限于:(i) 多数可用评估数据在训练阶段已被最强大模型接触过,(ii) 现有基准无法捕捉证明过程在多种维度上的任意复杂性。本文提出一个名为 MathGAP 的数据生成框架,用于评估 LLM 在具有任意复杂算术证明的问题上的表现。该框架根据对算术证明结构的指定,生成问题陈述与链式思维推理轨迹,支持对从简单到复杂的系统性泛化研究。使用 MathGAP 发现,随着证明树变深、变宽,模型性能显著下降;在复杂非线性证明结构中,即使最先进的模型也面临挑战。模型对句子顺序的微小变化敏感,但仍能解决部分复杂问题,表明推理泛化具有噪声性。
原文摘要 · Abstract (English)
Large language models (LLMs) can solve arithmetic word problems with high accuracy, but little is known about how well they generalize to more complex problems. This is difficult to study, as (i) much of the available evaluation data has already been seen by the most capable models during training, and (ii) existing benchmarks do not capture how problem proofs may be arbitrarily complex in various ways. In this paper, we present a data-generation framework for evaluating LLMs on problems with arbitrarily complex arithmetic proofs, called MathGAP. MathGAP generates problem statements and chain-of-thought reasoning traces according to specifications about their arithmetic proof structure, enabling systematic studies on easy-to-hard generalization with respect to complexity of proof trees. Using MathGAP, we find that LLMs show a significant decrease in performance as proofs get deeper and wider. This effect is more pronounced in complex, nonlinear proof structures, which are challenging even for the most capable models. The models are also sensitive to simple changes in sentence ordering. However, they remain capable of solving some complex problems, suggesting that reasoning generalization is noisy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。