首个面向统计推理的大型基准,专测模型在课程与科研级任务中的逻辑推导能力。
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
- 构建多智能体管道,将学术文献转为可评估的定理级推理题
- 覆盖超10万道题,含2万+基础题和8万+科研级证明题
- 提出自适应评分机制,精准衡量复杂推导过程,适合研究者与教育者
尽管大语言模型(LLMs)进展迅速,但现有评测基准对统计推理的覆盖仍不足,难以反映真实统计实践中层层递进、以证明为核心的特性。为此,我们提出首个涵盖课程与科研级场景的大规模统计推理基准——StatEval。该基准包含超过10万道精心筛选的问题,其中2万余道为基础性题目,覆盖本科与研究生课程内容;8万余道为从顶级统计期刊中提取的科研级证明任务。为构建此基准,我们开发了多智能体管道TRACE(Topology and Reasoning-Aware Context Extractor),结合人机协同验证,将非结构化学术文本转化为自洽的定理级推理任务。同时,我们提出自适应过程式评分流水线,实现对复杂统计证明的细粒度评估,超越仅比对最终答案的传统方式。实验表明,尽管大语言模型在基础任务上表现尚可,但在严谨的科研级推理中仍存在明显短板。此外,检索增强生成与领域对齐方法能持续提升模型性能。综上,StatEval不仅是一个评测基准,更成为推动大语言模型统计推理能力发展的基础设施。
原文摘要 · Abstract (English)
Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-driven nature of real statistical practice. To address this gap, we introduce \textbf{StatEval}, the first large-scale benchmark for statistical reasoning across curricular and research-level settings. StatEval includes over 100,000 curated problems, with 20,000+ foundational questions spanning undergraduate and graduate curricula and 80,000+ research-level proof tasks extracted from leading statistical journals. To construct StatEval, we develop \textbf{TRACE} (Topology and Reasoning-Aware Context Extractor), a multi-agent pipeline with human-in-the-loop validation that converts unstructured academic texts into self-contained theorem-level reasoning tasks. We also propose an Adaptive Process-Based Scoring Pipeline for complex statistical proofs, enabling fine-grained evaluation beyond final-answer matching. Experiments show that while LLMs perform reasonably on foundational tasks, they struggle with rigorous research-level reasoning. Beyond evaluation, StatEval serves as a resource for improving reasoning, as retrieval-augmented generation and domain-specific alignment consistently enhance performance. Together, these results establish StatEval as both a benchmark and an infrastructure for advancing statistical reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。