arXiv:2506.17363cs.CYcs.AI2025-06ACL被引 5

在477名研究生中测试大模型助教,发现其能有效支持编程教学。

A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant

  • 基于大模型构建虚拟助教,在真实课程中部署并持续收集反馈。
  • 分析3869次互动,识别出常见问题类型与学生参与模式。
  • 对比真人教师互动,揭示助教在教学中的潜力与局限,适合教育科技研究者参考。

由大型语言模型(LLMs)驱动的虚拟教学助理(VTAs)有望通过即时反馈和多轮交互提升学生学习效果。然而,现有针对其在真实课堂中的有效性与接受度的实证研究有限,实际影响尚不明确。本研究开发了一个基于LLM的虚拟助教,并将其部署于包含477名研究生的入门级人工智能编程课程中。为评估学生对助教表现认知随时间的变化,我们在课程不同阶段进行了三轮全面调查。同时,我们分析了3,869对师生-虚拟助教交互,识别出常见问题类型与参与模式,并与传统师生互动进行对比,以评估虚拟助教在学习过程中的作用。通过大规模实证研究与交互分析,我们评估了虚拟助教在真实课堂中部署的可行性,并指出了广泛推广的关键挑战。最后,我们公开了虚拟助教系统的源代码,以促进人工智能驱动教育的发展: exttt{https://github.com/sean0042/VTA}。

原文摘要 · Abstract (English)

Virtual Teaching Assistants (VTAs) powered by Large Language Models (LLMs) have the potential to enhance student learning by providing instant feedback and facilitating multi-turn interactions. However, empirical studies on their effectiveness and acceptance in real-world classrooms are limited, leaving their practical impact uncertain. In this study, we develop an LLM-based VTA and deploy it in an introductory AI programming course with 477 graduate students. To assess how student perceptions of the VTA's performance evolve over time, we conduct three rounds of comprehensive surveys at different stages of the course. Additionally, we analyze 3,869 student--VTA interaction pairs to identify common question types and engagement patterns. We then compare these interactions with traditional student--human instructor interactions to evaluate the VTA's role in the learning process. Through a large-scale empirical study and interaction analysis, we assess the feasibility of deploying VTAs in real-world classrooms and identify key challenges for broader adoption. Finally, we release the source code of our VTA system, fostering future advancements in AI-driven education: \texttt{https://github.com/sean0042/VTA}.

虚拟助教大模型教育AI实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。