提出可量化编码质量的计算方法,助力人机协作质性分析
A Computational Method for Measuring "Open Codes" in Qualitative Analysis
- 用大模型增强算法合并编码本,生成统一基准
- 设计覆盖、重叠、新颖性、差异四项指标评估编码质量
- 能识别冗余或虚假编码,适合研究者与AI协同使用
质性分析在社会科学中至关重要,其核心方法是归纳编码——研究者从数据本身提取并解释编码。然而,这种探索性方法难以满足对深度与多样性等方法论要求,尤其在越来越多研究者借助生成式AI辅助时更为突出。基于真实标签的度量方式因违背归纳编码的探索本质而不可靠,手动评估又耗时费力。本文提出一种理论驱动的计算方法,用于衡量人类与生成式AI的归纳编码结果。该方法首先利用大模型增强算法合并个体编码本;再通过四项新指标——覆盖率、重叠率、新颖性与差异性,评估每位编码者的贡献。在在线对话数据集上的两项实验表明:1)合并算法影响各项指标表现;2)指标在多次运行及不同大模型下具有稳定性和鲁棒性;3)可有效诊断编码问题,如过度编码或虚构(幻觉)编码。本工作为确保人机协同质性分析的方法严谨性提供了可靠路径。
原文摘要 · Abstract (English)
Qualitative analysis is critical to understanding human datasets in many social science disciplines. A central method in this process is inductive coding, where researchers identify and interpret codes directly from the datasets themselves. Yet, this exploratory approach poses challenges for meeting methodological expectations (such as ``depth'' and ``variation''), especially as researchers increasingly adopt Generative AI (GAI) for support. Ground-truth-based metrics are insufficient because they contradict the exploratory nature of inductive coding, while manual evaluation can be labor-intensive. This paper presents a theory-informed computational method for measuring inductive coding results from humans and GAI. Our method first merges individual codebooks using an LLM-enriched algorithm. It measures each coder's contribution against the merged result using four novel metrics: Coverage, Overlap, Novelty, and Divergence. Through two experiments on a human-coded online conversation dataset, we 1) reveal the merging algorithm's impact on metrics; 2) validate the metrics' stability and robustness across multiple runs and different LLMs; and 3) showcase the metrics' ability to diagnose coding issues, such as excessive or irrelevant (hallucinated) codes. Our work provides a reliable pathway for ensuring methodological rigor in human-AI qualitative analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。