arXiv:2410.04698cs.CL2024-10被引 15

评测大模型在长文本中的数学推理能力,发现顶级模型仍有明显短板。

MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs

  • 设计自动化基准测试,融合信息检索与复杂数学推理。
  • 顶级模型在128K上下文仅达51.26%准确率,表现不佳。
  • 适合关注长文本数学推理能力的研究者和开发者。

近期大型语言模型(LLMs)在长上下文场景中展现出多样化能力。尽管已有部分基准用于评估模型的长上下文能力,但缺乏专门评估模型在长文本中数学推理能力的基准,而这一能力对真实应用场景至关重要。本文提出MathHay,一个自动化基准,用于评估LLMs在长上下文中的数学推理能力。与以往聚焦于长文本中信息检索的基准(如Needle in a Haystack)不同,MathHay要求模型同时具备信息搜索与复杂数学推理能力。我们在八个顶尖的LLMs上进行了广泛实验,结果显示,即使是最优模型Gemini-1.5-Pro-002,在128K token长度下也仅达到51.26%的准确率,表明该任务仍存在巨大提升空间。

原文摘要 · Abstract (English)

Recent large language models (LLMs) have demonstrated versatile capabilities in long-context scenarios. Although some recent benchmarks have been developed to evaluate the long-context capabilities of LLMs, there is a lack of benchmarks evaluating the mathematical reasoning abilities of LLMs over long contexts, which is crucial for LLMs' application in real-world scenarios. In this paper, we introduce MathHay, an automated benchmark designed to assess the long-context mathematical reasoning capabilities of LLMs. Unlike previous benchmarks like Needle in a Haystack, which focus primarily on information retrieval within long texts, MathHay demands models with both information-seeking and complex mathematical reasoning abilities. We conduct extensive experiments on MathHay to assess the long-context mathematical reasoning abilities of eight top-performing LLMs. Even the best-performing model, Gemini-1.5-Pro-002, still struggles with mathematical reasoning over long contexts, achieving only 51.26% accuracy at 128K tokens. This highlights the significant room for improvement on the MathHay benchmark.

数学推理长文本基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。