arXiv:2603.26673cs.CYcs.AI2026-03

对比三款AI在教学中的表现,发现它们对不同教学法反应不一。

Can AI be a Teaching Partner? Evaluating ChatGPT, Gemini, and DeepSeek across Three Teaching Strategies

  • 用三种教学法测试ChatGPT、Gemini和DeepSeek的教学能力
  • 在提问引导法中,模型表现受提示词影响更大,效果差异明显
  • Gemini和ChatGPT整体得分更高,适合教学场景使用

大型语言模型(LLMs)被寄予厚望,可为学生提供解释、反馈与指导以支持学习。然而,尽管其迅速普及,关于其教学能力的实证研究仍有限。本文比较了ChatGPT、DeepSeek和Gemini作为教学代理的表现,采用三种教学策略:举例、解释与类比、苏格拉底式提问。六名人类评委在教授初学者C语言的背景下进行评估。结果表明,模型在举例和解释与类比策略中表现出相似互动模式;而在苏格拉底式提问中,模型对教学策略和初始提示更为敏感。总体而言,ChatGPT和Gemini得分较高,而DeepSeek得分较低,显示出不同模型间教学表现存在差异。

原文摘要 · Abstract (English)

There are growing promises that Large Language Models (LLMs) can support students' learning by providing explanations, feedback, and guidance. However, despite their rapid adoption and widespread attention, there is still limited empirical evidence regarding the pedagogical skills of LLMs. This article presents a comparative study of popular LLMs, namely, ChatGPT, DeepSeek, and Gemini, acting as teaching agents. An evaluation protocol was developed, focusing on three pedagogical strategies: Examples, Explanations and Analogies, and the Socratic Method. Six human judges conducted the evaluations in the context of teaching the C programming language to beginners. The results indicate that LLM models exhibited similar interaction patterns in the pedagogical strategies of Examples and Explanations and Analogies. In contrast, for the Socratic Method, the models showed greater sensitivity to the pedagogical strategy and the initial prompt. Overall, ChatGPT and Gemini received higher scores, whereas DeepSeek obtained lower scores across the criteria, indicating differences in pedagogical performance across models.

AI教学语言模型教育评估苏格拉底法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。