arXiv:2506.05990cs.SEcs.AI2025-06中稿 · BEA 2025被引 4

用AI自动生成编程竞赛测试用例,提升评测质量与效率

Leveraging Generative AI for Enhancing Automated Assessment in Programming Education Contests

  • 基于大语言模型生成高质量测试用例,自动化替代人工设计
  • 在5年级奥赛题中发现67%题目存在此前未察觉的错误
  • 适合教育机构和竞赛组织者快速提升评测能力

编程竞赛对培养学习者的计算思维与算法能力至关重要,但设计全面有效的测试用例对教师而言仍耗时且困难。本文提出一种基于自然语言处理的创新方法,利用生成式AI(大语言模型)自动创建高质量的编程竞赛测试用例。我们在多个数据集上进行了广泛评估,包括近25年罗马尼亚信息学奥林匹克竞赛(OJI)5年级组数据、Kilonova.ro平台近期赛事数据,以及国际团队信息学奥林匹克竞赛(IIOT)。结果表明,AI生成的测试用例显著提升了评估效果,成功识别出67%的OJI 5年级编程题中此前未被发现的错误。这凸显了该技术在形成性评估中的互补价值。我们开源了提示词、翻译后的数据集及方法,为教育工作者和竞赛组织者提供可直接集成的NLP工具,以提升评测质量、降低工作负担,并深化对学习者表现的理解。

原文摘要 · Abstract (English)

Competitive programming contests play a crucial role in cultivating computational thinking and algorithmic skills among learners. However, generating comprehensive test cases to effectively assess programming solutions remains resource-intensive and challenging for educators. This paper introduces an innovative NLP-driven method leveraging generative AI (large language models) to automate the creation of high-quality test cases for competitive programming assessments. We extensively evaluated our approach on diverse datasets, including 25 years of Romanian Informatics Olympiad (OJI) data for 5th graders, recent competitions hosted on the Kilonova.ro platform, and the International Informatics Olympiad in Teams (IIOT). Our results demonstrate that AI-generated test cases substantially enhanced assessments, notably identifying previously undetected errors in 67% of the OJI 5th grade programming problems. These improvements underscore the complementary educational value of our technique in formative assessment contexts. By openly sharing our prompts, translated datasets, and methodologies, we offer practical NLP-based tools that educators and contest organizers can readily integrate to enhance assessment quality, reduce workload, and deepen insights into learner performance.

编程教育生成式AI自动评测竞赛优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。