arXiv:2608.03038cs.CLstat.AP2026-08

评估大模型统计推理能力,不只看答对率,还分析解释思路和风格。

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

论文配图:Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 融合准确率、回答行为与文本分析,多维度评估模型推理能力。
  • 模型准确率在55%至78%之间波动,差距显著。
  • 同厂商模型解释风格更相似,说明存在厂商特有表达习惯。

统计推理具有多维性,但现有大语言模型(LLMs)评估多聚焦于答案正确率,忽视模型如何构建和传达统计解释。本研究通过结合响应准确率、响应行为、结构主题建模及词汇相似性分析,建立多维评估框架。该框架应用于15个当前主流大模型对90道来自高中、本科及研究生水平四场统计考试题目的回答。模型准确率在55%至78%间波动,差异显著;结构主题建模显示所有模型均呈现相似的统计推理概念结构;词汇相似性分析则揭示厂商间存在微弱但稳定的解释风格差异,同一厂商模型生成的解释比跨厂商模型更相似。结果表明,仅以准确率衡量大模型统计推理能力是片面的,结合响应行为与解释内容的互补分析,可更全面评估生成式AI的统计推理表现。

原文摘要 · Abstract (English)

Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\% to 78\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.

统计推理多维评估语言模型解释分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。