提出无限复杂度数学题生成器,揭示大模型长文本推理的瓶颈
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
- 基于计算图抽象生成可无限扩展难度的数学题
- 发现推理性能随复杂度呈饱和式下降,计算量翻倍仅带来线性提升
- 适合研究长文本推理、模型可扩展性的学者和工程师
长上下文大语言模型在信息检索和长文档问答中表现强劲,但面对前沿数学研究等复杂推理任务时仍显不足。现有评估基准缺乏定量基础。受GSM-8K问题可抽象为计算图的启发,我们开发了一种数学题生成器,可通过添加冗余节点与边引入噪声,实现对难度与上下文长度的精细控制,生成无限复杂度的算术题。基于此构建了GSM-Infinite基准,对现有大模型进行全面评估。结果表明,随着复杂度增加,推理性能呈现一致的S型衰减;且推理计算量呈指数增长时,性能仅线性提升。这些发现揭示了当前长上下文大模型在推理能力上的根本局限。GSM-Infinite提供了一个可扩展、可控的测试平台,可用于系统研究并推动大模型在复杂长上下文中的推理能力发展。
原文摘要 · Abstract (English)
Long-context large language models (LLMs) have recently shown strong performance in information retrieval and long-document QA. However, to tackle the most challenging intellectual problems, LLMs must reason effectively in long and complex contexts (e.g., frontier mathematical research). Studying how LLMs handle increasing reasoning complexity and context length is essential, yet existing benchmarks lack a solid basis for quantitative evaluation. Inspired by the abstraction of GSM-8K problems as computational graphs, and the ability to introduce noise by adding unnecessary nodes and edges, we develop a grade school math problem generator capable of producing arithmetic problems with infinite difficulty and context length under fine-grained control. Using our newly synthesized GSM-Infinite benchmark, we comprehensively evaluate existing LLMs. We find a consistent sigmoid decline in reasoning performance as complexity increases, along with a systematic inference scaling trend: exponentially increasing inference computation yields only linear performance gains. These findings underscore the fundamental limitations of current long-context LLMs and the key challenges in scaling reasoning capabilities. Our GSM-Infinite benchmark provides a scalable and controllable testbed for systematically studying and advancing LLM reasoning in long and complex contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。