arXiv:2601.14479cs.CL2026-01被引 3

测试大模型统计推理能力,发现微调后可媲美本科生水平。

Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks

  • 用自研数据集微调开源大模型,提升其统计推理能力。
  • 微调后模型在复杂统计任务上达到统计专业学生水平。
  • 模型能自我评估答案质量,适合教育平台与科研验证。

本文研究大语言模型(LLMs)解决统计任务的能力及其对推理质量的评估能力。尽管当前最先进的LLMs在自然语言处理任务中表现优异,但其在中等复杂度统计问题上的表现仍不明确。我们使用自建数据集对部分开源LLMs进行微调,以增强其统计推理能力,并将其性能与人类基准评分进行对比。结果表明,微调后的模型在高级统计任务上的表现已接近统计学专业学生水平。微调效果具有架构依赖性,部分模型取得显著提升,显示出在教育科技和统计分析辅助系统中的部署潜力。此外,我们发现LLMs自身在判断答案质量(包括解释与推理)方面远优于传统指标(如BLEU或BertScore)。这一自评估能力为统计教育平台的自动化测评及自动化分析工具的质量保障提供了可扩展方案。潜在应用还包括学术与产业界的研究方法验证、数据分析工作流的质量控制。

原文摘要 · Abstract (English)

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of NLP tasks, their competence in addressing even moderately complex statistical challenges is not well understood. We have fine-tuned selected open-source LLMs on a specially developed dataset to enhance their statistical reasoning capabilities, and compared their performance with the human scores used as a benchmark. Our results show that the fine-tuned models achieve better performance on advanced statistical tasks on the level comparable to a statistics student. Fine-tuning demonstrates architecture-dependent improvements, with some models showing significant performance gains, indicating clear potential for deployment in educational technology and statistical analysis assistance systems. We also show that LLMs themselves can be far better judges of the answers quality (including explanation and reasoning assessment) in comparison to traditional metrics, such as BLEU or BertScore. This self-evaluation capability enables scalable automated assessment for statistical education platforms and quality assurance in automated analysis tools. Potential applications also include validation tools for research methodology in academic and industry settings, and quality control mechanisms for data analysis workflows.

大模型统计推理自评估教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。