arXiv:2504.13882cs.HCcs.CL2025-04中稿 · the Workshop on "F…被引 1

用大模型自动评估教学对话中的五种关键策略效果。

Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation

  • 用GPT-3.5+少样本提示,判断教学对话中策略是否恰当使用。
  • 识别正确策略的召回率在0.327~0.432之间,误判率较低。
  • 对“帮助学生应对不平等”策略识别效果最好,适合教育AI研究者参考。

本研究提出一种基于大语言模型(LLMs)的自动化系统,用于评估五类关键教学策略的有效性:1. 给予有效表扬,2. 响应错误,3. 判断学生已知内容,4. 帮助学生应对不平等,5. 回应消极自我评价。利用来自教师-学生聊天室语料库(Teacher-Student Chatroom Corpus)的公开数据集,系统将每项策略分类为“按预期使用”或“未按预期使用”。研究采用GPT-3.5结合少样本提示进行策略评估与对话分析。结果显示,五项策略的真负率(TNR)在0.655至0.738之间,召回率(Recall)在0.327至0.432之间,表明模型能较好排除错误分类,但对正确策略的识别能力仍有不足。其中,“帮助学生应对不平等”策略表现最佳,TNR为0.738,召回率为0.432。研究揭示了大模型在教学策略分析中的潜力,并指出未来可借助更先进模型实现更细致反馈。

原文摘要 · Abstract (English)

Our study introduces an automated system leveraging large language models (LLMs) to assess the effectiveness of five key tutoring strategies: 1. giving effective praise, 2. reacting to errors, 3. determining what students know, 4. helping students manage inequity, and 5. responding to negative self-talk. Using a public dataset from the Teacher-Student Chatroom Corpus, our system classifies each tutoring strategy as either being employed as desired or undesired. Our study utilizes GPT-3.5 with few-shot prompting to assess the use of these strategies and analyze tutoring dialogues. The results show that for the five tutoring strategies, True Negative Rates (TNR) range from 0.655 to 0.738, and Recall ranges from 0.327 to 0.432, indicating that the model is effective at excluding incorrect classifications but struggles to consistently identify the correct strategy. The strategy \textit{helping students manage inequity} showed the highest performance with a TNR of 0.738 and Recall of 0.432. The study highlights the potential of LLMs in tutoring strategy analysis and outlines directions for future improvements, including incorporating more advanced models for more nuanced feedback.

教学评估大模型应用对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。