arXiv:2409.18596cs.AIcs.CL2024-09中稿 · SIGCSE-Virtual 202…被引 8

构建跨学科短答评分基准,评估自动评分系统表现

ASAG2024: A Combined Benchmark for Short Answer Grading

  • 整合七个数据集形成统一评分框架
  • 大模型方法达新高分但仍远低于人类水平
  • 为教育AI提供可比测试基准,适合教育技术研究者

开放性问题比封闭性问题更能检验深入理解,是更优的评估方式,但人工批改费时且易受主观偏见影响。为此,学术界致力于自动化评分。短答评分(SAG)系统旨在自动打分。尽管SAG方法不断发展,目前尚无涵盖不同学科、评分尺度和分布的综合性基准,难以评估现有方法的泛化能力。本文初步提出ASAG2024联合基准,将七个常用短答评分数据集统一为相同结构与评分标准。我们评估了若干近期SAG方法,发现基于大语言模型的方法虽达新高分,仍远未接近人类表现。这为未来人机协同评分系统研究开辟了方向。

原文摘要 · Abstract (English)

Open-ended questions test a more thorough understanding than closed-ended questions and are often a preferred assessment method. However, open-ended questions are tedious to grade and subject to personal bias. Therefore, there have been efforts to speed up the grading process through automation. Short Answer Grading (SAG) systems aim to automatically score students' answers. Despite growth in SAG methods and capabilities, there exists no comprehensive short-answer grading benchmark across different subjects, grading scales, and distributions. Thus, it is hard to assess the capabilities of current automated grading methods in terms of their generalizability. In this preliminary work, we introduce the combined ASAG2024 benchmark to facilitate the comparison of automated grading systems. Combining seven commonly used short-answer grading datasets in a common structure and grading scale. For our benchmark, we evaluate a set of recent SAG methods, revealing that while LLM-based approaches reach new high scores, they still are far from reaching human performance. This opens up avenues for future research on human-machine SAG systems.

自动评分教育AI大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。