构建多语言多模态编程竞赛数据集,评测大模型在真实教育场景中的表现。
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
- 设计中英双语、图文混合的编程选择题数据集,模拟真实竞赛场景。
- 发现大模型在纸上推理题上表现优于代码生成题,且英语题表现略好于罗马尼亚语。
- 开源数据集与学习应用,助力多语言编程教育研究与实践。
大型语言模型(LLMs)的快速发展已深刻影响计算机科学(CS)教育领域,其在代码任务和问题求解中展现出卓越能力,引发对其在高级编程场景中潜力与局限性的关注。本研究提出一个新型双语(英语-罗马尼亚语)多模态(文本与图像)的多项选择题数据集,源自高水平计算机科学竞赛。该数据集的特殊之处在于:部分题目更适合纸上演算推理,而另一些则更宜通过编写代码解决。我们系统评估了当前主流大模型在该数据集上的表现,分析其在理论编程任务中的性能。结果揭示了现有模型的优势与不足,包括语言选择(英语与罗马尼亚语)的影响,为模型在编程教育及竞赛环境中的适用性提供了洞见。同时,本文还探讨了教育公平性与评估完整性等关键伦理问题,旨在指导未来教育实践与政策制定。为促进后续研究,数据集将公开提供中英文版本。此外,我们还发布了一个专为罗马尼亚学生设计的教育应用,支持基于该数据集的互动式自测与练习。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving, raising questions about their potential and limitations in advanced CS contexts. This study presents a novel bilingual (English-Romanian) multimodal (text and image) dataset of multiple-choice questions derived from a high-level computer science competition. A particularity of our dataset is that the problems are conceived such that some of them are easier solved using reasoning on paper, while for others writing code is more efficient. We systematically evaluate State of The Art LLMs on this dataset, analyzing their performance on theoretical programming tasks. Our findings reveal the strengths and limitations of current LLMs, including the influence of language choice (English vs. Romanian), providing insights into their applicability in CS education and competition settings. We also address critical ethical considerations surrounding educational integrity and the fairness of assessments in the context of LLM usage. These discussions aim to inform future educational practices and policies. To support further research, our dataset will be made publicly available in both English and Romanian. Additionally, we release an educational application tailored for Romanian students, enabling them to self-assess using the dataset in an interactive and practice-oriented environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。