arXiv:2409.15112cs.AI2024-09

测试ChatGPT在西班牙语编程考试中的解题与评分能力

ChatGPT as a Solver and Grader of Programming Exams written in Spanish

  • 用ChatGPT解决真实本科计算机课程的西班牙语编程题
  • 仅能有效处理简单编码任务,复杂问题表现差
  • 适合研究AI在教育评估中的局限性,尤其语言非英文场景

评估大型语言模型(LLMs)在辅助教学任务中的能力正受到越来越多关注。本文评估了ChatGPT在解答由正规计算机科学本科课程提供的西班牙语编程考试题时的表现。结果表明,该AI模型仅对简单编码任务有效,对于复杂问题的求解或对他人代码的评估能力远未达到实用水平。作为研究的一部分,我们还发布了一个新的编程题目语料库及对应的解题与评分提示,可供其他研究团队进一步使用。

原文摘要 · Abstract (English)

Evaluating the capabilities of Large Language Models (LLMs) to assist teachers and students in educational tasks is receiving increasing attention. In this paper, we assess ChatGPT's capacities to solve and grade real programming exams, from an accredited BSc degree in Computer Science, written in Spanish. Our findings suggest that this AI model is only effective for solving simple coding tasks. Its proficiency in tackling complex problems or evaluating solutions authored by others are far from effective. As part of this research, we also release a new corpus of programming tasks and the corresponding prompts for solving the problems or grading the solutions. This resource can be further exploited by other research teams.

编程评估大模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。