用生成式AI自动分类导师对话行为,准确率达80%
Automated Classification of Tutors' Dialogue Acts Using Generative AI: A Case Study Using the CIMA Corpus
- 用GPT-4+定制提示词实现对话行为自动标注
- 达到80%准确率,F1-score 0.81,与人工标注高度一致
- 适合教育对话分析研究者快速处理标注数据
本研究探讨生成式AI在自动化分类导师对话行为(Dialogue Acts, DAs)中的应用,旨在减少传统人工标注的时间与精力消耗。案例研究使用开源CIMA语料库,其中导师回应已预先标注为四类对话行为。测试了GPT-3.5-turbo与GPT-4模型,采用定制化提示词。结果表明,GPT-4取得80%准确率、加权F1-score 0.81、Cohen's Kappa 0.74,优于基线表现,显示出与人工标注的高度一致性。研究结果表明,生成式AI在教育对话分析中具备高效且可及的潜力。同时强调任务特定标签定义与上下文信息对自动化标注质量的重要性,并指出生成式AI使用中的伦理问题,呼吁负责任透明的研究实践。研究脚本公开于https://github.com/liqunhe27/Generative-AI-for-educational-dialogue-act-tagging。
原文摘要 · Abstract (English)
This study explores the use of generative AI for automating the classification of tutors' Dialogue Acts (DAs), aiming to reduce the time and effort required by traditional manual coding. This case study uses the open-source CIMA corpus, in which tutors' responses are pre-annotated into four DA categories. Both GPT-3.5-turbo and GPT-4 models were tested using tailored prompts. Results show that GPT-4 achieved 80% accuracy, a weighted F1-score of 0.81, and a Cohen's Kappa of 0.74, surpassing baseline performance and indicating substantial agreement with human annotations. These findings suggest that generative AI has strong potential to provide an efficient and accessible approach to DA classification, with meaningful implications for educational dialogue analysis. The study also highlights the importance of task-specific label definitions and contextual information in enhancing the quality of automated annotation. Finally, it underscores the ethical considerations associated with the use of generative AI and the need for responsible and transparent research practices. The script of this research is publicly available at https://github.com/liqunhe27/Generative-AI-for-educational-dialogue-act-tagging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。