arXiv:2501.10421cs.CYcs.AI2025-01被引 14

用大模型自动评编程作业,反馈更一致、更有建设性。

CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback

  • 用思维链提示增强大模型推理,对齐人工评分标准。
  • 集成多个大模型并做一致性测试,提升评分准确率。
  • 适合教育科技、自动化评测系统开发者参考。

批改编程作业对提升学生编程能力与代码风格至关重要。本文提出自动化评分框架CodEv,利用大语言模型(LLMs)生成一致且具建设性的反馈。通过引入思维链(Chain of Thought, CoT)提示技术,增强模型推理能力,确保评分与人工评价一致。框架还采用大模型集成策略,并结合一致性验证测试,以提升评分准确性与可靠性。实验表明,该框架使用较小规模的LLM即可达到与人工评分相当的效果。对LLM的评估与一致性测试进一步验证了其生成分数与反馈的可靠性。

原文摘要 · Abstract (English)

Grading programming assignments is crucial for guiding students to improve their programming skills and coding styles. This study presents an automated grading framework, CodEv, which leverages Large Language Models (LLMs) to provide consistent and constructive feedback. We incorporate Chain of Thought (CoT) prompting techniques to enhance the reasoning capabilities of LLMs and ensure that the grading is aligned with human evaluation. Our framework also integrates LLM ensembles to improve the accuracy and consistency of scores, along with agreement tests to deliver reliable feedback and code review comments. The results demonstrate that the framework can yield grading results comparable to human evaluators, by using smaller LLMs. Evaluation and consistency tests of the LLMs further validate our approach, confirming the reliability of the generated scores and feedback.

编程评测大模型应用教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。