arXiv:2502.17657stat.APcs.AI2025-02被引 13

构建首个用于评估大模型统计编程能力的开源数据集

StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis

  • 设计包含任务、代码与人工评分的三元数据集
  • 涵盖多种分析类型,测试主流大模型生成SAS代码表现
  • 适合评估大模型在数据科学中的实用性和改进方向

大语言模型的编码能力为机器学习和数据科学中的自动统计分析带来了新机遇。但在广泛应用前,亟需评估其生成代码的准确性。当前主要障碍是缺乏针对统计代码(如SAS和R)的基准数据集。为此,本文提出StatLLM,一个开源数据集,用于评估大模型在统计分析中的表现。该数据集包含三个核心部分:涵盖多种分析类型与数据集的统计分析任务,以及对应任务下由ChatGPT 3.5、ChatGPT 4.0和Llama 3.1生成的SAS代码;由领域专家对生成代码的正确性、有效性、可读性、可执行性和输出准确性进行评分的人工评价结果。该数据集还可用于评估和优化自然语言处理指标、提升大模型在统计编程中的性能,并推动下一代统计软件的研发,对数据科学与机器学习研究具有重要意义。

原文摘要 · Abstract (English)

The coding capabilities of large language models (LLMs) have opened up new opportunities for automatic statistical analysis in machine learning and data science. However, before their widespread adoption, it is crucial to assess the accuracy of code generated by LLMs. A major challenge in this evaluation lies in the absence of a benchmark dataset for statistical code (e.g., SAS and R). To fill in this gap, this paper introduces StatLLM, an open-source dataset for evaluating the performance of LLMs in statistical analysis. The StatLLM dataset comprises three key components: statistical analysis tasks, LLM-generated SAS code, and human evaluation scores. The first component includes statistical analysis tasks spanning a variety of analyses and datasets, providing problem descriptions, dataset details, and human-verified SAS code. The second component features SAS code generated by ChatGPT 3.5, ChatGPT 4.0, and Llama 3.1 for those tasks. The third component contains evaluation scores from human experts in assessing the correctness, effectiveness, readability, executability, and output accuracy of the LLM-generated code. We also illustrate the unique potential of the established benchmark dataset for (1) evaluating and enhancing natural language processing metrics, (2) assessing and improving LLM performance in statistical coding, and (3) developing and testing of next-generation statistical software - advancements that are crucial for data science and machine learning research.

大模型评测统计分析SAS代码数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。