arXiv:2506.12066cs.CL2025-06被引 1

用教材生成问答并自动评分,减轻师生负担

Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading

  • 按文档视觉布局切分文本,提升下游任务精度
  • 基于教材生成高质量问题与参考答案
  • 验证大模型可泛化到短答案自动评分

数字技术在教育中日益普及,用于减轻师生负担。然而,生成开放性习题及评分仍耗时费力。本论文提出一个基于课堂材料生成问题并自动评分的系统框架。针对PDF文档,提出一种结合视觉布局的文本切块方法,显著提升检索增强生成(RAG)等下游任务的准确性。实验证明,可从学习材料中生成高质量问题与参考答案。同时构建了一个新的短答案自动评分基准,用于比较不同评分系统。评估显示,大语言模型(LLMs)能通过预训练任务泛化至短答案评分,模型参数量越大,性能越优。目前系统仍需人工审核,尤其在考试场景中。

原文摘要 · Abstract (English)

Digital technologies are increasingly used in education to reduce the workload of teachers and students. However, creating open-ended study or examination questions and grading their answers is still a tedious task. This thesis presents the foundation for a system that generates questions grounded in class materials and automatically grades student answers. It introduces a sophisticated method for chunking documents with a visual layout, specifically targeting PDF documents. This method enhances the accuracy of downstream tasks, including Retrieval Augmented Generation (RAG). Our thesis demonstrates that high-quality questions and reference answers can be generated from study material. Further, it introduces a new benchmark for automated grading of short answers to facilitate comparison of automated grading systems. An evaluation of various grading systems is conducted and indicates that Large Language Models (LLMs) can generalise to the task of automated grading of short answers from their pre-training tasks. As with other tasks, increasing the parameter size of the LLMs leads to greater performance. Currently, available systems still need human oversight, especially in examination scenarios.

自动评分问答生成大模型应用教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。