arXiv:2504.05239cs.CL2025-04中稿 · IEEE TALE 2025被引 28

用人类参与的交互式框架,让大模型更精准地自动批改开放式答题。

LLM-based Automated Grading with Human-in-the-Loop

  • 通过人机协作让大模型主动提问,动态优化评分标准。
  • 在评分标准评估中显著提升准确率,逼近人工水平。
  • 适合需要高精度、可解释性评分的教育场景使用。

人工智能技术,尤其是大语言模型(LLMs),正推动教育领域的变革。在自动短答案评分(ASAG)任务中,基于LLM的方法不仅超越了传统方法,还能实现基于评分量规的复杂评估,而不仅仅是对比预设的‘黄金答案’。然而,现有全自动方案在量规评分中仍难以达到人类水平。本文提出一种人机协同(HITL)框架GradeHITL,利用大模型的生成能力向人类专家提出问题,结合其反馈动态优化评分标准。该自适应机制显著提升了评分准确性,优于现有方法,使自动评分更接近人工水平。

原文摘要 · Abstract (English)

The rise of artificial intelligence (AI) technologies, particularly large language models (LLMs), has brought significant advancements to the field of education. Among various applications, automatic short answer grading (ASAG), which focuses on evaluating open-ended textual responses, has seen remarkable progress with the introduction of LLMs. These models not only enhance grading performance compared to traditional ASAG approaches but also move beyond simple comparisons with predefined "golden" answers, enabling more sophisticated grading scenarios, such as rubric-based evaluation. However, existing LLM-powered methods still face challenges in achieving human-level grading performance in rubric-based assessments due to their reliance on fully automated approaches. In this work, we explore the potential of LLMs in ASAG tasks by leveraging their interactive capabilities through a human-in-the-loop (HITL) approach. Our proposed framework, GradeHITL, utilizes the generative properties of LLMs to pose questions to human experts, incorporating their insights to refine grading rubrics dynamically. This adaptive process significantly improves grading accuracy, outperforming existing methods and bringing ASAG closer to human-level evaluation.

自动评分人机协同大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。