评测大模型统计分析能力,发现其仍远未达到可靠水平。
StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs

- 构建双组件基准:封闭题库覆盖18类统计主题,开放题模拟真实建模任务。
- 顶尖模型在封闭题仅达68.6%准确率,开放题平均得分61.86。
- 揭示大模型在工具使用、方法选择和全流程建模上仍有明显短板。
统计分析是需要领域知识与工具熟练度的复杂领域。尽管已有研究评估大语言模型(LLMs)在此领域的表现,但现有基准在范围和形式上仍显局限。为此,我们提出StatABench(统计分析基准),系统评估LLMs的统计分析能力。该基准包含两个互补部分:Stat-Closed包含404道题,覆盖18个统计主题,题型包括单选、填空、决策和实际应用;Stat-Open包含30个来自专业竞赛的复杂开放式建模任务。我们使用LangChain MCP框架和多个数据科学代理评估模型,并通过验证过的LLM-as-Judge协议评估开放题解法。实验显示,即使GPT-5.1在Stat-Closed上也仅达68.6%准确率,最佳开源模型为60.6%;在Stat-Open上,最优代理框架平均得分为61.86。结果表明当前大模型在可靠统计分析方面仍存在显著差距,暴露出工具引导推理、方法决策和端到端建模等持续挑战。
原文摘要 · Abstract (English)
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, existing benchmarks remain limited in scope and format. To bridge this gap, we introduce StatABench (Statistical AnalysisBenchmark), a benchmark designed to systematically assess LLMs' statistical analysis capabilities. StatABench comprises two complementary components: Stat-Closed, containing 404 questions across 18 statistical topics in multiple formats (multiple-choice, fill-in-the-blank, decision-making, and practical application), and Stat-Open, featuring 30 complex open-ended modeling tasks adapted from professional competitions. We evaluate diverse LLMs using the LangChain MCP framework and multiple data science agents, and assess Stat-Open solutions via a validated LLM-as-Judge protocol. Experiments show that even GPT-5.1 achieves only 68.6% on Stat-Closed, while the best open-source model reaches 60.6%. On Stat-Open, the top agent framework scores 61.86 on average. These results reveal the gap between current LLMs and reliable statistical analysis, highlighting persistent challenges in tool-grounded reasoning, methodological decision-making, and end-to-end statistical modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。