用大模型辅助教育AI对话分析,人类主导概念构建与验证。
Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
- 人机分阶段协作:机器生成标签,人类定义概念和合并类别。
- 从4.5万条对话中提炼出72个代码,涵盖19类6域,可靠度经三名专家验证。
- 适合需高质量编码的教育研究者,强调人类在解释中的决定性作用。
随着质性研究数据规模扩大,人工编码已难应对,大语言模型(LLMs)常被提议作为分析助手。本文详述一个三阶段人机协作流程,将开放式、轴心式与选择性编码应用于45,000条中小学教师与生成式AI平台的交互消息,构建层级化代码本。在各阶段,LLMs大规模生成候选标签与结构化标注,而人类研究人员始终掌握类别定义、合并决策与解释框架的最终裁量权。随后,三位具备教育领域专长的编码员对2,560条独立样本进行系统编码,通过多标签一致率测量迭代校准,达成信度,并补充了5个原流程未识别的新代码。最终代码本包含72项,分属19个类别与6个领域。文章反思了该流程中的关键方法选择,包括分析单元设定、将LLM视为标注工具而非解释主体、多标签编码下的信度测量方式,以及人类专业知识在关键节点的不可替代性。该过程提供可审计模板,供质性研究者在保留人类解释权威的前提下,安全使用大模型辅助代码本开发。
原文摘要 · Abstract (English)
Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educational domain expertise applied the codebook to an independent sample of 2,560 messages, established reliability through iterative calibration using set-valued agreement measures appropriate for multi-label annotation, and extended the instrument with five codes that the LLM-assisted phases had not surfaced. The final codebook comprises 72 items within 19 categories and six domains. We reflect on the methodological decisions the pipeline required, including the choice of a conversational unit of analysis, the treatment of the LLM as a labeling instrument rather than an interpretive agent, the measurement of intercoder agreement under multi-label coding, and the conditions under which human domain expertise remained decisive. The account is offered as an auditable template for qualitative researchers considering LLM assistance in codebook development while preserving human interpretive authority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。