大模型数学辅导接近专家水平,但教学风格和语言特点有差异。
Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles
- 对比专家、新手与7个大小不一的LLM,分析辅导策略与语言特征。
- 大模型平均表现接近专家,但更冗长、更礼貌、词汇更丰富。
- 强调推理与重述比高情商表达更影响教学效果,适合教育研究者参考。
近期研究探索了大语言模型(LLMs)生成数学辅导回应的应用,但其教学行为与人类专家的契合度尚不明确。我们分析了一个数学补救对话数据集,其中专家导师、新手导师及七个不同规模的LLM(包含开源与商业模型)针对相同学生错误作出回应。研究考察了辅导回应中的教学策略与语言特征,包括应答(重述与复述)、强调准确性和推理、词汇多样性、可读性、礼貌程度与代理度。结果表明,专家导师的表现优于新手,大模型普遍得分高于小模型,平均接近专家水平。然而,大模型在教学风格上存在系统性差异:较少使用专家常用的讨论性策略,却产生更长、词汇更丰富、更礼貌的回应。回归分析显示,强调准确性和推理、重述与复述、词汇多样性与感知教学质量正相关,而高代理度与礼貌语言则呈负相关。研究强调,在评估辅导回应时需综合分析教学策略与语言特征,适用于人类导师与智能辅导系统的比较研究。
原文摘要 · Abstract (English)
Recent work has explored the use of large language models (LLMs) to generate tutoring responses in mathematics, yet it remains unclear how closely their instructional behavior aligns with expert human practice. We analyze a dataset of math remediation dialogues in which expert tutors, novice tutors, and seven LLMs of varying sizes, comprising both open-weight and commercial models, respond to the same student errors. We examine instructional strategies and linguistic characteristics of tutoring responses, including uptake (restating and revoicing), pressing for accuracy and reasoning, lexical diversity, readability, politeness, and agency. We find that expert tutors produce higher-quality responses than novices, and that larger LLMs generally receive higher pedagogical quality ratings than smaller models, approaching expert performance on average. However, LLMs exhibit systematic differences in their instructional profiles: they underuse discursive strategies characteristic of expert tutors while generating longer, more lexically diverse, and more polite responses. Regression analyses show that pressing for accuracy and reasoning, restating and revoicing, and lexical diversity, are positively associated with perceived pedagogical quality, whereas higher levels of agentic and polite language are negatively associated. These findings highlight the importance of analyzing instructional strategies and linguistic characteristics when evaluating tutoring responses across human tutors and intelligent tutoring systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。