评测大模型解罗马尼亚编程竞赛题表现,发现不同模型差距明显
Evaluating the Performance of Large Language Models in Competitive Programming: A Multi-Year, Multi-Grade Analysis
- 用2002-2023年304道真题测试大模型代码能力
- GPT-4在中学级题目上表现最优,准确率超70%
- 开源模型如CodeLlama也展现实用潜力,适合教学辅助
本研究评估大语言模型(LLMs)在解决罗马尼亚信息学奥赛县赛级别编程题的表现。罗马尼亚是计算机竞赛强国,其高标准赛事为评估提供了理想环境。研究收集并分析了2002至2023年间共304道题目,重点关注模型使用C++和Python编写的解决方案。通过多轮尝试与反馈机制,评估了包括GPT-4在内的多个闭源模型及CodeLlama、RoMistral等开源模型。结果表明,模型表现随题目难度与年级显著变化。其中,GPT-4在中学组题目中表现优异,准确率达70%以上,展现出作为教学辅助工具的潜力。同时,不同模型在代码质量与风格上存在明显差异。
原文摘要 · Abstract (English)
This study explores the performance of large language models (LLMs) in solving competitive programming problems from the Romanian Informatics Olympiad at the county level. Romania, a leading nation in computer science competitions, provides an ideal environment for evaluating LLM capabilities due to its rich history and stringent competition standards. We collected and analyzed a dataset comprising 304 challenges from 2002 to 2023, focusing on solutions written by LLMs in C++ and Python for these problems. Our primary goal is to understand why LLMs perform well or poorly on different tasks. We evaluated various models, including closed-source models like GPT-4 and open-weight models such as CodeLlama and RoMistral, using a standardized process involving multiple attempts and feedback rounds. The analysis revealed significant variations in LLM performance across different grades and problem types. Notably, GPT-4 showed strong performance, indicating its potential use as an educational tool for middle school students. We also observed differences in code quality and style across various LLMs
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。