arXiv:2504.03877cs.LG2025-04被引 5

用概念评分标准提升大模型的学业评估与数据生成能力

Concept-based Rubrics Improve LLM Formative Assessment and Data Synthesis

  • 基于学科概念设计评分标准,显著提升大模型评估表现
  • 在多类数据集上实现接近小模型的评估准确率
  • 可生成高质量合成数据,用于训练轻量高效监督模型

STEM领域形成性评估旨在通过识别学生当前理解水平来促进后续学习。现有研究表明,生成式大语言模型(LLMs)在开放问答中的构造应答评估性能远低于依赖高质量标注数据的监督分类器。本文证明,基于概念的评分标准能显著提升LLM表现,缩小其与需大量训练数据的小型监督模型之间的差距。在概念评分标准使LLM表现良好的数据集上,我们进一步验证其可生成高质量合成数据,用于训练轻量级、高性能的监督模型。实验覆盖多种包含不同质量标签的STEM学生作答数据集,包括一个含人工智能辅助回答的真实世界数据集,带来新的评估考量。

原文摘要 · Abstract (English)

Formative assessment in STEM topics aims to promote student learning by identifying students' current understanding, thus targeting how to promote further learning. Previous studies suggest that the assessment performance of current generative large language models (LLMs) on constructed responses to open-ended questions is significantly lower than that of supervised classifiers trained on high-quality labeled data. However, we demonstrate that concept-based rubrics can significantly enhance LLM performance, which narrows the gap between LLMs as off-the shelf assessment tools, and smaller supervised models, which need large amounts of training data. For datasets where concept-based rubrics allow LLMs to achieve strong performance, we show that the concept-based rubrics help the same LLMs generate high quality synthetic data for training lightweight, high-performance supervised models. Our experiments span diverse STEM student response datasets with labels of varying quality, including a new real-world dataset that contains some AI-assisted responses, which introduces additional considerations.

大模型评估教育科技数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。