构建细粒度评分基准,提升大模型对简答的评估能力
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
- 设计分步评分机制,支持精细化打分与错误类型标注
- 涵盖1030道真实考题和4109份学生作答,专家标注
- 揭示大模型在科学类题目上的评分难题,验证少样本提示有效
主观答题评分(SAG)在教育、标准化测试和自动化评估中至关重要,尤其针对短答案评分(SAS)。现有方法常给出粗粒度分数,缺乏详细推理。尽管大语言模型(LLMs)展现出零样本评估潜力,但仍存在偏见、与人类判断不一致及评分透明度不足等问题。为此,我们提出SAS-Bench,一个专为基于LLM的SAS任务设计的基准。该基准提供细粒度、分步评分,包含专家标注的错误类别,并整合来自真实学科考试的多样化题型。它支持对模型推理过程和可解释性的深入评估。我们还开源了包含1,030个问题和4,109个学生回答的数据集,均由领域专家标注。通过多种LLM的全面实验,我们识别出科学类题目评分中的主要挑战,并验证了少样本提示在提升评分准确性方面的有效性。本研究为构建更鲁棒、公平且具教育意义的LLM评估系统提供了关键洞见。
原文摘要 · Abstract (English)
Subjective Answer Grading (SAG) plays a crucial role in education, standardized testing, and automated assessment systems, particularly for evaluating short-form responses in Short Answer Scoring (SAS). However, existing approaches often produce coarse-grained scores and lack detailed reasoning. Although large language models (LLMs) have demonstrated potential as zero-shot evaluators, they remain susceptible to bias, inconsistencies with human judgment, and limited transparency in scoring decisions. To overcome these limitations, we introduce SAS-Bench, a benchmark specifically designed for LLM-based SAS tasks. SAS-Bench provides fine-grained, step-wise scoring, expert-annotated error categories, and a diverse range of question types derived from real-world subject-specific exams. This benchmark facilitates detailed evaluation of model reasoning processes and explainability. We also release an open-source dataset containing 1,030 questions and 4,109 student responses, each annotated by domain experts. Furthermore, we conduct comprehensive experiments with various LLMs, identifying major challenges in scoring science-related questions and highlighting the effectiveness of few-shot prompting in improving scoring accuracy. Our work offers valuable insights into the development of more robust, fair, and educationally meaningful LLM-based evaluation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。