测试大模型解定积分能力,发现差距明显且难度影响准确率
INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
- 构建含符号与数值真值的定积分评测集
- 九个主流大模型准确率差异显著,难度越高错误越多
- 适合关注数学推理能力评估的研究者使用
我们提出INTEGRALBENCH,一个专注于评估大语言模型(LLM)在定积分问题上表现的基准。该基准提供符号与数值两种真值解,并由人工标注题目难度。对九个最先进的LLM进行评估,结果揭示了显著的性能差距,且题目难度与模型准确率存在强相关性,为这一挑战性领域建立了基线指标。INTEGRALBENCH旨在通过专门针对定积分计算的严谨评估框架,推动自动化数学推理的发展。
原文摘要 · Abstract (English)
We present INTEGRALBENCH, a focused benchmark designed to evaluate Large Language Model (LLM) performance on definite integral problems. INTEGRALBENCH provides both symbolic and numerical ground truth solutions with manual difficulty annotations. Our evaluation of nine state-of-the-art LLMs reveals significant performance gaps and strong correlations between problem difficulty and model accuracy, establishing baseline metrics for this challenging domain. INTEGRALBENCH aims to advance automated mathematical reasoning by providing a rigorous evaluation framework specifically tailored for definite integral computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。