arXiv:2410.01240cs.CLcs.HC2024-10被引 9

用大模型自动完成话语分析编码,提升效率。

Automatic deductive coding in discourse analysis: an application of large language models in learning analytics

  • 用提示工程引导大模型实现自动编码
  • 在小样本下准确率优于传统方法和BERT
  • 适合教育数据分析研究者快速处理话语数据

演绎编码是学习科学与学习分析中常用的话语分析方法,需研究人员根据理论框架手动标注所有话语,耗时费力。大语言模型(如GPT)的出现为自动化演绎编码提供了新路径。本文比较了三种基于不同AI技术的分类方法:传统文本分类(特征工程)、BERT类预训练模型和GPT类大语言模型(LLM)。在两个不同数据集上评估其性能,结果表明,在训练样本有限的情况下,采用提示工程的GPT方法在准确率和Kappa值上均优于其他两种方法。通过设计详细提示结构,本研究展示了大语言模型在自动演绎编码中的应用潜力。

原文摘要 · Abstract (English)

Deductive coding is a common discourse analysis method widely used by learning science and learning analytics researchers for understanding teaching and learning interactions. It often requires researchers to manually label all discourses to be analyzed according to a theoretically guided coding scheme, which is time-consuming and labor-intensive. The emergence of large language models such as GPT has opened a new avenue for automatic deductive coding to overcome the limitations of traditional deductive coding. To evaluate the usefulness of large language models in automatic deductive coding, we employed three different classification methods driven by different artificial intelligence technologies, including the traditional text classification method with text feature engineering, BERT-like pretrained language model and GPT-like pretrained large language model (LLM). We applied these methods to two different datasets and explored the potential of GPT and prompt engineering in automatic deductive coding. By analyzing and comparing the accuracy and Kappa values of these three classification methods, we found that GPT with prompt engineering outperformed the other two methods on both datasets with limited number of training samples. By providing detailed prompt structures, the reported work demonstrated how large language models can be used in the implementation of automatic deductive coding.

话语分析大模型学习分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。