评测大模型在长文本中的数学推理能力,发现顶级模型仍有明显短板。
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
- 设计自动化基准测试,融合信息检索与复杂数学推理。
- 顶级模型在128K上下文仅达51.26%准确率,表现不佳。
- 适合关注长文本数学推理能力的研究者和开发者。
近期大型语言模型(LLMs)在长上下文场景中展现出多样化能力。尽管已有部分基准用于评估模型的长上下文能力,但缺乏专门评估模型在长文本中数学推理能力的基准,而这一能力对真实应用场景至关重要。本文提出MathHay,一个自动化基准,用于评估LLMs在长上下文中的数学推理能力。与以往聚焦于长文本中信息检索的基准(如Needle in a Haystack)不同,MathHay要求模型同时具备信息搜索与复杂数学推理能力。我们在八个顶尖的LLMs上进行了广泛实验,结果显示,即使是最优模型Gemini-1.5-Pro-002,在128K token长度下也仅达到51.26%的准确率,表明该任务仍存在巨大提升空间。
原文摘要 · Abstract (English)
Recent large language models (LLMs) have demonstrated versatile capabilities in long-context scenarios. Although some recent benchmarks have been developed to evaluate the long-context capabilities of LLMs, there is a lack of benchmarks evaluating the mathematical reasoning abilities of LLMs over long contexts, which is crucial for LLMs' application in real-world scenarios. In this paper, we introduce MathHay, an automated benchmark designed to assess the long-context mathematical reasoning capabilities of LLMs. Unlike previous benchmarks like Needle in a Haystack, which focus primarily on information retrieval within long texts, MathHay demands models with both information-seeking and complex mathematical reasoning abilities. We conduct extensive experiments on MathHay to assess the long-context mathematical reasoning abilities of eight top-performing LLMs. Even the best-performing model, Gemini-1.5-Pro-002, still struggles with mathematical reasoning over long contexts, achieving only 51.26% accuracy at 128K tokens. This highlights the significant room for improvement on the MathHay benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。