学生问LLM的多数是操作类问题,尤其在备考时更明显。
"How Do I ...?": Procedural Questions Predominate Student-LLM Chatbot Conversations
- 分析6113条对话,发现操作类问题占主导,备考时尤为突出。
- 用LLM做分类评分一致性高于人类,但现有分类框架覆盖不足。
- 建议结合对话分析法,深入理解多轮互动中的学习行为。
基于大语言模型(LLM)的教育聊天机器人在提供学习支持方面具有潜力,但也存在风险与挑战。当学生遇到障碍时,会通过提出受阻驱动的问题来寻求帮助,这类问题直接影响用户输入和聊天机器人的教学效果。本研究基于两个不同学习情境的数据集——形成性自学与总结性评估作业,分析了6,113条消息,采用11种不同的LLM和3名人类评估者,使用四种现有分类框架对学生的提问进行标注。结果表明,使用LLM作为评估者可达到中等到良好的一致性,且优于人类评估者。数据发现,在两种学习情境中,'操作类'问题均占主导地位,且在备考准备阶段更为显著。该结果为利用LLM进行学生提问分类提供了基础,但同时也揭示了现有分类框架的局限性:其结构过于简单,难以涵盖复合提示语义的丰富性,仅能提供对聊天机器人整合风险与收益的片面理解。未来建议采用对话分析方法,如话语心理学中的分析策略,以捕捉对话的细微层次与多轮互动特征。
原文摘要 · Abstract (English)
Providing scaffolding through educational chatbots built on Large Language Models (LLM) has potential risks and benefits that remain an open area of research. When students navigate impasses, they ask for help by formulating impasse-driven questions. Within interactions with LLM chatbots, such questions shape the user prompts and drive the pedagogical effectiveness of the chatbot's response. This paper focuses on such student questions from two datasets of distinct learning contexts: formative self-study, and summative assessed coursework. We analysed 6,113 messages from both learning contexts, using 11 different LLMs and three human raters to classify student questions using four existing schemas. On the feasibility of using LLMs as raters, results showed moderate-to-good inter-rater reliability, with higher consistency than human raters. The data showed that 'procedural' questions predominated in both learning contexts, but more so when students prepare for summative assessment. These results provide a basis on which to use LLMs for classification of student questions. However, we identify clear limitations in both the ability to classify with schemas and the value of doing so: schemas are limited and thus struggle to accommodate the semantic richness of composite prompts, offering only partial understanding the wider risks and benefits of chatbot integration. In the future, we recommend an analysis approach that captures the nuanced, multi-turn nature of conversation, for example, by applying methods from conversation analysis in discursive psychology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。