用概念图分析多模态作答,评估学生思维质量。
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
- 以概念图为框架,从多模态回答中推断学生思维质量。
- 最佳模型准确率仅40%,误差1.1单位,未达人类水平。
- 可辅助教师高效诊断学情,定制教学方案。
科学、技术、工程和数学(STEM)中的认知模型在评估学生对主题的概念理解方面起着关键作用。它们不仅揭示学生知道什么,还反映其在不同情境中应用、关联和整合概念的能力。因此,学生的回答是其理解质量的重要指标,不应仅被评分。然而,从学生作答中推断这些认知模型具有挑战性,因为它需要深度推理能力。我们提出了MMGrader,一种利用概念图作为分析框架,从学生的多模态响应中推断其认知模型质量的方法。在对9个公开可用模型的评估中,我们发现表现最好的模型仍未能达到人类水平:准确率约为40%,预测误差为1.1单位,且评分分布与人类评分模式基本一致。若能进一步提升准确率,这些模型可成为教师的有效助手,帮助其高效推断整个班级的认知模型,从而设计有针对性的辅导和讲座,增强学生集体表现较弱的知识领域。
原文摘要 · Abstract (English)
STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts across various contexts. Thus, students' responses are critical markers of the quality of their understanding and not entities that should be merely graded. However, inferring these mental models from student answers is challenging as it requires deep reasoning skills. We propose MMGrader, an approach that infers the quality of students' mental models from their multimodal responses using concept graphs as an analytical framework. In our evaluation with 9 openly available models, we found that the best-performing models fall short of human-level performance. This is because they only achieved an accuracy of approximately 40%, a prediction error of 1.1 units, and a scoring distribution fairly aligned with human scoring patterns. With improved accuracy, these can be highly effective assistants to teachers in inferring the mental models of their entire classrooms, enabling them to do so efficiently and help improve their pedagogies more effectively by designing targeted help sessions and lectures that strengthen areas where students collectively demonstrate lower proficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。