arXiv:2602.01015cs.CLcs.CY2026-02被引 2

LLM在模拟学习者思考时过于流畅自信,高估了学生表现。

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

  • 用真实学生思考录音对比LLM生成的推理过程
  • 模型推理更连贯冗长,错误率被严重高估
  • 适合研究AI辅导系统可信度的学者参考

大型语言模型(LLMs)正越来越多地用于基于AI的辅导系统。它们能否忠实模拟新手的学习思维和元认知判断?现有评估侧重解题准确率,忽视了人类学习中碎片化、不完美的推理特征。本研究基于630条多步骤化学辅导问题中的学生思考录音,结合学生提示使用、尝试次数和问题上下文的日志数据,比较了在最小与扩展上下文提示下,LLM生成的推理与真人学习者言语的差异,并评估模型预测学生每一步成功的能力。尽管GPT-4.1能生成流畅且语境恰当的延续,其推理却系统性地表现出过度连贯、冗长且变异程度低于真人思考。这种现象在提示中引入更丰富的问题背景时加剧。模型对学习者表现的判断持续偏高。这些发现揭示了用LLM模拟学习过程的元认知局限,归因于训练数据中缺乏情感表达及工作记忆限制等真实学习特征。本研究提出的评估框架可指导未来更真实支持新手学习与自我调节的自适应系统设计。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive judgments? Existing evaluations emphasize problem-solving accuracy, overlooking the fragmented and imperfect reasoning that characterizes human learning. We evaluate LLMs as novices using 630 think-aloud utterances from multi-step chemistry tutoring problems with problem-solving logs of student hint use, attempts, and problem context. We compare LLM-generated reasoning to human learner utterances under minimal and extended contextual prompting, and assess the models' ability to predict step-level learner success. Although GPT-4.1 generates fluent and contextually appropriate continuations, its reasoning is systematically over-coherent, verbose, and less variable than human think-alouds. These effects intensify with a richer problem-solving context during prompting. Learner performance was consistently overestimated. These findings highlight epistemic limitations of simulating learning with LLMs. We attribute these limitations to LLM training data, including expert-like solutions devoid of expressions of affect and working memory constraints during problem solving. Our evaluation framework can guide future design of adaptive systems that more faithfully support novice learning and self-regulation using generative artificial intelligence.

大模型评估学习模拟认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。