最新大模型在算法考试中表现媲美顶尖学生,可生成高质量教学内容。
From Struggle (06-2024) to Mastery (02-2025) LLMs Conquer Advanced Algorithm Exams and Pave the Way for Editorial Generation
- 在罗马上大学和英文版算法考题上测试多款大模型表现。
- 最新模型得分接近顶尖学生,多步推理能力强,但图论题仍存难点。
- 可自动生成教学评注,助力教师提升反馈质量,适合教育场景使用。
本文全面评估了当前先进大语言模型在具有挑战性的大学级算法考试中的表现。通过在罗马尼亚试题及其高质量英文翻译上测试多个模型,分析其解题能力、一致性及多语言表现。实证研究表明,最新模型不仅得分接近顶尖学生,且在复杂多步骤算法问题上展现出稳健的推理能力,尽管在图相关任务上仍有困难。基于此,我们探索了大模型在教育环境中生成高质量教学评注的潜力,为教师提供增强学生反馈的强大工具。文中提出的洞见与最佳实践,为生成式AI在高级算法教育中的进一步融合铺平道路。
原文摘要 · Abstract (English)
This paper presents a comprehensive evaluation of the performance of state-of-the-art Large Language Models (LLMs) on challenging university-level algorithms exams. By testing multiple models on both a Romanian exam and its high-quality English translation, we analyze LLMs' problem-solving capabilities, consistency, and multilingual performance. Our empirical study reveals that the most recent models not only achieve scores comparable to top-performing students but also demonstrate robust reasoning skills on complex, multi-step algorithmic challenges, even though difficulties remain with graph-based tasks. Building on these findings, we explore the potential of LLMs to support educational environments through the generation of high-quality editorial content, offering instructors a powerful tool to enhance student feedback. The insights and best practices discussed herein pave the way for further integration of generative AI in advanced algorithm education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。