arXiv:2502.13117stat.APcs.AI2025-02被引 6

评估大模型生成统计代码的准确性与可靠性。

Performance Evaluation of Large Language Models in Statistical Programming

  • 对比ChatGPT、Llama生成SAS代码的语法正确性。
  • 模型在复杂任务中准确率不足,易产生冗余或错误结果。
  • 适合研究人员和数据分析师参考大模型在统计编程中的局限。

大型语言模型(LLMs)的编程能力已革新自动代码生成,并为自动化统计分析开辟新路径。然而,在广泛采用前,其生成代码的有效性与质量需系统评估。尽管影响力日益增长,针对统计代码生成的全面评估仍较缺乏。本文评估了包括两个版本ChatGPT和一个版本Llama在内的大模型在SAS编程领域的表现。研究涵盖多种统计主题与数据集的分析任务,每项任务包含问题描述、数据信息及人工验证的SAS代码。通过专家评审对生成代码的正确性、有效性、可读性、可执行性及输出结果准确性进行综合评估。评分分析显示,虽然大模型能生成语法正确的代码,但在需要深层领域理解的任务中表现不佳,常产生冗余或错误结果。本研究为大模型在统计编程中的能力与局限提供了重要洞见,为未来智能编码系统的发展提供指导。

原文摘要 · Abstract (English)

The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.

统计编程大模型SAS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。