用大模型自动分析对话数据,提升编码准确率。
LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding
- 基于对话行为与事件设计分步提示,增强上下文理解
- 多模型协作预测,事件编码准确率显著提升
- 通过事件间关联一致性校验,适合教育对话研究者
对话数据是理解学习过程的重要来源,揭示学生在协作讨论中的互动方式及其对知识建构的影响。大型语言模型(LLMs)为对话数据的自动化编码带来了新机遇,但对话固有的上下文复杂性给模型理解带来挑战。本研究提出一种新型的LLM辅助对话数据自动编码方法,其创新点在于:1)基于对话特定特征——交流行为与交流事件,采用角色提示与思维链方法分别生成代码;2)联合使用GPT-4-turbo、GPT-4o和DeepSeek进行多模型协作预测;3)利用事件与行为间的关联性,通过GPT-4o实现一致性校验。实验表明,该方法显著提升了编码准确性,且行为预测准确率始终高于事件预测。本研究为提升对话数据编码精度提供了新方法框架,并提供了一种可扩展的解决方案,以应对对话分析中的上下文难题。
原文摘要 · Abstract (English)
Dialogue data has been a key source for understanding learning processes, offering critical insights into how students engage in collaborative discussions and how these interactions shape their knowledge construction. The advent of Large Language Models (LLMs) has introduced promising opportunities for advancing qualitative research, particularly in the automated coding of dialogue data. However, the inherent contextual complexity of dialogue presents unique challenges for these models, especially in understanding and interpreting complex contextual information. This study addresses these challenges by developing a novel LLM-assisted automated coding approach for dialogue data. The novelty of our proposed framework is threefold: 1) We predict the code for an utterance based on dialogue-specific characteristics -- communicative acts and communicative events -- using separate prompts following the role prompts and chain-of-thoughts methods; 2) We engaged multiple LLMs including GPT-4-turbo, GPT-4o, DeepSeek in collaborative code prediction; 3) We leveraged the interrelation between events and acts to implement consistency checking using GPT-4o. In particular, our contextual consistency checking provided a substantial accuracy improvement. We also found the accuracy of act predictions was consistently higher than that of event predictions. This study contributes a new methodological framework for enhancing the precision of automated coding of dialogue data as well as offers a scalable solution for addressing the contextual challenges inherent in dialogue analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。