arXiv:2603.10233cs.CL2026-03中稿 · , https://sgrades

构建统一评估框架,测试AI在作文与问答评分中的泛化能力

S-GRADES -- Studying Generalization of Student Response Assessments in Diverse Evaluative Settings

  • 整合14个数据集的统一评测平台,支持跨任务评估
  • 发现主流大模型在不同题型间表现差异显著,泛化能力有限
  • 适合教育AI研究者、评测工具开发者使用

自动作文评分(AES)关注逻辑连贯性与论证质量,自动短答案评分(ASAG)侧重事实正确性与概念理解。尽管目标一致,两类研究长期孤立发展,存在数据集分散、评价指标不一、社区割裂等问题。本文提出S-GRADES(Studying Generalization of Student Response Assessments in Diverse Evaluative Settings),一个基于网页的基准评测平台,将14个多样化的评分数据集统一接入,提供标准化访问和可复现的评估协议。该平台完全开源且可扩展,支持持续集成新数据集与评估场景。我们以三种先进大语言模型为基础,采用多种提示策略进行评测,并分析示例选择与跨数据集示例迁移的影响。结果揭示了基准评测能有效暴露不同任务间可靠性与泛化能力的差距,凸显了跨范式标准化评估的重要性。

原文摘要 · Abstract (English)

Evaluating student responses, from long essays to short factual answers, is a key challenge in educational NLP. Automated Essay Scoring (AES) focuses on holistic writing qualities such as coherence and argumentation, while Automatic Short Answer Grading (ASAG) emphasizes factual correctness and conceptual understanding. Despite their shared goal, these paradigms have progressed in isolation with fragmented datasets, inconsistent metrics, and separate communities. We introduce S-GRADES (Studying Generalization of Student Response Assessments in Diverse Evaluative Settings), a web-based benchmark that consolidates 14 diverse grading datasets under a unified interface with standardized access and reproducible evaluation protocols. The benchmark is fully open-source and designed for extensibility, enabling continuous integration of new datasets and evaluation settings. To demonstrate the utility of S-GRADES, we evaluate three state-of-the-art large language models across the benchmark using multiple reasoning strategies in prompting. We further examine the effects of exemplar selection and cross-dataset exemplar transfer. Our analyses illustrate how benchmark-driven evaluation reveals reliability and generalization gaps across essay and short-answer grading tasks, highlighting the importance of standardized, cross-paradigm assessment.

教育AI评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。