对比元学习与视觉大模型自动评分手写图表,发现不同任务下各有优劣。
Automated Grading of Students' Handwritten Graphs: A Comparison of Meta-Learning and Vision-Large Language Models
- 用多模态元学习模型处理手写图+文本的自动评分
- 元学习在二分类任务中优于视觉大模型,三分类中后者略胜一筹
- 适合教育技术研究者和自动化评估系统开发者参考
随着在线学习兴起,数学作业高效且一致的评分需求在过去十年显著上升。机器学习(ML),特别是自然语言处理(NLP),已被广泛用于自动评分以文本和/或数学表达为主的回答。然而,尽管手写图表在科学、技术、工程和数学(STEM)课程中普遍使用,针对此类作答的自动评分研究仍十分有限。本研究实现了用于评分包含学生手写图表和文本的图像的多模态元学习模型,并进一步将这些模型与视觉大语言模型(VLLMs)的性能进行比较。在本机构真实数据集上的评估结果表明,在二分类任务中,表现最佳的元学习模型优于VLLMs;而在更复杂的三分类任务中,表现最佳的VLLMs则略胜一筹。尽管VLLMs展现出良好前景,其可靠性和实际应用价值仍不确定,需进一步研究。
原文摘要 · Abstract (English)
With the rise of online learning, the demand for efficient and consistent assessment in mathematics has significantly increased over the past decade. Machine Learning (ML), particularly Natural Language Processing (NLP), has been widely used for autograding student responses, particularly those involving text and/or mathematical expressions. However, there has been limited research on autograding responses involving students' handwritten graphs, despite their prevalence in Science, Technology, Engineering, and Mathematics (STEM) curricula. In this study, we implement multimodal meta-learning models for autograding images containing students' handwritten graphs and text. We further compare the performance of Vision Large Language Models (VLLMs) with these specially trained metalearning models. Our results, evaluated on a real-world dataset collected from our institution, show that the best-performing meta-learning models outperform VLLMs in 2-way classification tasks. In contrast, in more complex 3-way classification tasks, the best-performing VLLMs slightly outperform the meta-learning models. While VLLMs show promising results, their reliability and practical applicability remain uncertain and require further investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。