arXiv:2608.22993cs.CLcs.HC2026-08

研究大模型在辅导中提供帮助的程度,发现多数回答直接给解法。

LLM Pedagogical Behavior in AI Tutoring Interactions

  • 构建五级辅助程度量表,量化大模型辅导时的介入程度。
  • 95%以上回复属于解释或直接解题,帮助过强。
  • 辅导方式影响学生后续互动,但对考试成绩预测作用有限。

学生越来越多地将大模型用作课程学习和解题辅导工具,但对其在真实学习互动中提供的帮助程度了解甚少。这很重要,因为辅导响应在直接助人完成任务的程度上差异显著。本文将这一维度操作化为支架水平,并开发出一个五级量表,经人类标注验证后用于描述响应所提供的直接帮助程度。我们对203名大学生在大学人工智能课程中的14,637条大模型回应进行了分析。结果显示,绝大多数回应集中在高辅助水平,超过95%被归类为解释或直接求解。支架水平与学生后续对话行为系统相关,但除了先验成绩和对话行为外,对三场后续考试表现的预测能力几乎无增益。研究为大模型辅导中的辅助程度提供了实证基准,并建立了评估不同辅导设计改变该辅助水平的测量框架。

原文摘要 · Abstract (English)

Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.

大模型辅导教育人工智能支架教学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。