评测大模型诊断手写数学题认知能力,发现表现不佳且易误判模糊证据。
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
- 构建包含639份手写答题的数学认知诊断数据集MathCog
- 18个大模型在模糊证据下F1均低于0.5,误判率高
- 适合教育AI研发者关注证据感知与教师协同设计
学生的手写数学作答蕴含丰富的认知过程信息,远超最终答案本身。本文研究当前大语言模型(LLMs)在从此类作答中诊断认知技能的表现。由于学生答题常省略步骤或仅提供含糊、依赖上下文的线索,现有模型在此类复杂情境下的表现仍不明确。为此,我们构建了MathCog基准数据集,包含639份学生作答和110道数学题,共3,036条由教师标注的认知技能判断,依据TIMSS框架并附证据强度标签(明显/模糊)。评估18个主流LLM后发现:(1)所有模型性能均较差(F1 < 0.5),无论能力高低;(2)在模糊证据条件下性能急剧下降。错误分析显示,模型普遍存在将模糊证据误判为明显证据、过度解读微弱线索及虚构不存在证据的问题。研究对教育场景中基于大模型的认知诊断系统设计提出建议,强调证据感知与教师介入的重要性。
原文摘要 · Abstract (English)
Students' handwritten math work provides a rich resource for diagnosing cognitive skills, as it captures intermediate reasoning beyond final answers. We investigate how current large language models (LLMs) perform in diagnosing cognitive skills from such work. However, student responses vary widely, often omitting steps or providing only vague, contextually implicit evidence. Despite recent advances in LLMs' multimodal and reasoning capabilities, their performance under such conditions remains underexplored. To address this gap, we constructed MathCog, a benchmark dataset containing 3,036 diagnostic verdicts across 639 student responses to 110 math problems, annotated by teachers using TIMSS-grounded cognitive skill checklists with evidential strength labels (Evident/Vague). Evaluating 18 LLMs, we find that (1) all models underperform (F1 < 0.5) regardless of capability, and (2) performance degrades sharply under vague evidence. Error analysis reveals systematic patterns: models frequently misattribute Vague evidence as Evident, overthink minimal cues, and hallucinate nonexistent evidence. We discuss implications for evidence-aware, teacher-in-the-loop designs for LLM-based cognitive diagnosis in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。