用大模型分析医学生错题,找出隐藏的误解点。
What Don't You Understand? Using Large Language Models to Identify and Characterize Student Misconceptions About Challenging Topics
- 结合测验数据与大模型,识别难点知识点。
- 发现多个核心概念存在普遍误解,专家评价准确率高。
- 适合教育研究者和在线课程设计者使用。
本研究提出一种系统方法,通过量化表现分析与大语言模型(LLM)评估相结合,识别和刻画在线生物医学课程中学生的认知误区。分析了5门课程共9个学期的数据,涵盖3,802名医学生。每门课程包含40至50个主题聚焦的测验,首先基于首次作答的多选题(MCQ)表现,识别出在课程目标中居于核心地位且持续困难的主题。随后,利用生成式AI分析三类数据:测验题目内容、学生作答模式及授课讲义。该方法揭示了仅靠成绩数据无法察觉的认知误区,且专家对大模型识别出的误解质量评定为优秀。教师访谈也证实,基于数据的难点识别具有实用价值,并与教学观察一致。该方法可规模化应用于以测验为主的教学环境,为未来课程迭代提供精准干预路径,并可通过后续测验表现评估干预效果。
原文摘要 · Abstract (English)
This study presents a systematic approach to identifying and characterizing student misconceptions in online learning environments through a novel combination of quantitative performance analysis and large language model (LLM) assessment. We analyzed data from 9 course periods across 5 online biomedical science courses, encompassing 3,802 medical student enrollments. Using data from 40-50 topic-focused quizzes per course, we developed a two-stage methodology. First, we identified challenging central topics using quiz-level performance metrics. Second, we employed LLMs to characterize the underlying misconceptions in these high-priority areas. By examining student performance on first attempts across primarily multiple-choice questions (MCQs), we identified consistently challenging topics that were also central to course objectives. We then leveraged recent advances in generative AI to analyze three distinct data sources in combination: quiz question content, student response patterns, and lecture transcripts. This approach revealed actionable insights about student misconceptions that were not apparent from performance data alone. The quality of the LLM-identified misconceptions was rated as excellent by subject matter experts. We also conducted teacher interviews to assess the perceived utility of our topic identification method. Faculty found that data-driven identification of challenging topics was valuable and corroborated their own classroom observations. This methodology provides a scalable approach to characterizing student difficulties in learning environments where quizzes are used. Our findings demonstrate the potential for targeted and potentially personalized interventions in future course iterations, with clear pathways for measuring intervention effectiveness through follow-up quiz performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。