arXiv:2604.23347cs.CL2026-04

评测大模型在数据结构考试中的表现,揭示其真实水平。

Evaluating Large Language Models on Computer Science University Exams in Data Structures

论文配图:Evaluating Large Language Models on Computer Science University Exams in Data Structures
图 1 · 摘自论文原文
  • 构建特拉维夫大学数据结构考试新基准,含闭合与多选题。
  • GPT-4o与Claude 3.5表现最优,数学推理能力仍有限。
  • 适合教育研究者与大模型评估人员参考。

我们对大型语言模型(LLMs)在计算机科学(CS)数据结构考试题上的表现进行了全面评估。工作引入了一个新的基准数据集,包含特拉维夫大学(TAU)的考试题目,旨在评估大模型处理闭合题和多选题的能力。我们评估了OpenAI的GPT-4o和Anthropic的Claude 3.5,以及两个较小的模型Mathstral 7B和LLaMA 3 8B在TAU考试基准上的表现。研究结果揭示了当前大模型在计算机科学教育任务中的实际能力边界。

原文摘要 · Abstract (English)

We present a comprehensive evaluation of Large Language Models (LLMs) on Computer Science (CS) Data Structure examination questions. Our work introduces a new benchmark dataset comprising exam questions from Tel Aviv University (TAU), curated to assess LLMs' abilities in handling closed and multiple-choice questions. We evaluated the performance of OpenAI's GPT 4o and Anthropic's Claude 3.5, popular LLMs, alongside two smaller LLMs, Mathstral 7B and LLaMA 3 8B, across the TAU exams benchmark. Our findings provide insight into the current capabilities of LLMs in CS education.

大模型评估数据结构考试评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。