测试大模型在巴西计算机考研题上的表现,发现顶尖模型已超人类考生。
Assessing the Capability of LLMs in Solving POSCOMP Questions
- 用巴西计算机考研题评测多个大模型的答题能力
- ChatGPT-4在2022年答对57题,2023年超过所有人类考生
- 新模型持续超越人类,文本理解强但图像解析仍弱
大型语言模型(LLMs)在自然语言处理方面取得显著进展,但在计算机科学等专业领域的表现仍不明确。本文以巴西计算机学会(SBC)主办的权威研究生入学考试POSCOMP为基准,评估LLMs在该领域的实际能力。研究选取了ChatGPT-4、Gemini 1.0 Advanced、Claude 3 Sonnet和Le Chat Mistral Large四款模型,测试其在2022年和2023年两届考卷中的表现。结果显示,模型在文本类题目上表现优异,其中ChatGPT-4在2022年考卷中答对57/69题,领先于其他模型,并在2023年超过所有参加考试的人类考生。进一步测试2022–2024年考卷的更新模型(o1、Gemini 2.5 Pro、Claude 3.7 Sonnet、o3-mini-high),发现它们在各年份均持续超越人类平均与顶尖水平。尽管模型在图像理解任务上仍存在短板,但整体趋势表明,先进大模型在专业领域问题求解方面已具备较强竞争力。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have significantly expanded the capabilities of artificial intelligence in natural language processing tasks. Despite this progress, their performance in specialized domains such as computer science remains relatively unexplored. Understanding the proficiency of LLMs in these domains is critical for evaluating their practical utility and guiding future developments. The POSCOMP, a prestigious Brazilian examination used for graduate admissions in computer science promoted by the Brazlian Computer Society (SBC), provides a challenging benchmark. This study investigates whether LLMs can match or surpass human performance on the POSCOMP exam. Four LLMs - ChatGPT-4, Gemini 1.0 Advanced, Claude 3 Sonnet, and Le Chat Mistral Large - were initially evaluated on the 2022 and 2023 POSCOMP exams. The assessments measured the models' proficiency in handling complex questions typical of the exam. LLM performance was notably better on text-based questions than on image interpretation tasks. In the 2022 exam, ChatGPT-4 led with 57 correct answers out of 69 questions, followed by Gemini 1.0 Advanced (49), Le Chat Mistral (48), and Claude 3 Sonnet (44). Similar trends were observed in the 2023 exam. ChatGPT-4 achieved the highest performance, surpassing all students who took the POSCOMP 2023 exam. LLMs, particularly ChatGPT-4, show promise in text-based tasks on the POSCOMP exam, although image interpretation remains a challenge. Given the rapid evolution of LLMs, we expanded our analysis to include more recent models - o1, Gemini 2.5 Pro, Claude 3.7 Sonnet, and o3-mini-high - evaluated on the 2022-2024 POSCOMP exams. These newer models demonstrate further improvements and consistently surpass both the average and top-performing human participants across all three years.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。