用往年相似题目数据,让AI自动批改简答,准确率达人类水平。
AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
- 利用往年相近题目的学生作答数据训练AI评分模型。
- 在部分场景下比GPT-4等大模型高10%-20%准确率。
- 无需预设评分标准,适合实际教学环境快速部署。
开放性简答题对促进思维生成和检验核心概念理解至关重要,但教师时间有限、班级规模大等因素导致难以及时提供详细反馈。手动批改耗时,而现有自动化方法难以泛化到所有可能的回答场景。本文提出一种基于近域数据(如往年相似题目作答)的新型实用评分方法。该方法不依赖预设评分标准,可有效提升自动化评分准确性。实验表明,其性能显著优于当前最优机器学习模型及未微调的大语言模型(如GPT 3.5、GPT 4、GPT 4o),在某些情况下高出10%-20%,即便为这些模型提供了参考答案。研究还揭示了近域数据带来的准确率与数据量优势,是首个系统形式化近域数据用于自动化短答评分的工作。
原文摘要 · Abstract (English)
Constructed-response questions are crucial to encourage generative processing and test a learner's understanding of core concepts. However, the limited availability of instructor time, large class sizes, and other resource constraints pose significant challenges in providing timely and detailed evaluation, which is crucial for a holistic educational experience. In addition, providing timely and frequent assessments is challenging since manual grading is labor intensive, and automated grading is complex to generalize to every possible response scenario. This paper proposes a novel and practical approach to grade short-answer constructed-response questions. We discuss why this problem is challenging, define the nature of questions on which our method works, and finally propose a framework that instructors can use to evaluate their students' open-responses, utilizing near-domain data like data from similar questions administered in previous years. The proposed method outperforms the state of the art machine learning models as well as non-fine-tuned large language models like GPT 3.5, GPT 4, and GPT 4o by a considerable margin of over 10-20% in some cases, even after providing the LLMs with reference/model answers. Our framework does not require pre-written grading rubrics and is designed explicitly with practical classroom settings in mind. Our results also reveal exciting insights about learning from near-domain data, including what we term as accuracy and data advantages using human-labeled data, and we believe this is the first work to formalize the problem of automated short answer grading based on the near-domain data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。