对比人类与大模型在归纳编码中的表现,发现两者优势恰好相反。
Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
- 用大模型自动进行数据驱动的归纳编码,不预设标签
- 人类擅长复杂句编码,模型更准于简单句
- 模型标签更接近标准答案但常被专家打低分
本文研究定性数据分析的自动化,聚焦基于大语言模型(LLMs)的归纳编码。与依赖预定义标签的演绎方法不同,该研究关注从数据中自然涌现标签的归纳过程。评估了六种开源大模型的表现,并对比人类专家。专家对所编码引语的难度进行了主观评分。结果揭示出奇特的两极分化:人类在复杂句子上表现稳定,但在简单句子上出错较多;而大模型则恰恰相反。此外,研究通过与测试集的黄金标准对比,分析了人类和大模型标签的系统性偏差。尽管人类标注有时偏离标准答案,但常被其他人类评价为更优;部分大模型虽更贴近真实标签,却获得专家更低评分。
原文摘要 · Abstract (English)
This paper investigates the automation of qualitative data analysis, focusing on inductive coding using large language models (LLMs). Unlike traditional approaches that rely on deductive methods with predefined labels, this research investigates the inductive process where labels emerge from the data. The study evaluates the performance of six open-source LLMs compared to human experts. As part of the evaluation, experts rated the perceived difficulty of the quotes they coded. The results reveal a peculiar dichotomy: human coders consistently perform well when labeling complex sentences but struggle with simpler ones, while LLMs exhibit the opposite trend. Additionally, the study explores systematic deviations in both human and LLM generated labels by comparing them to the golden standard from the test set. While human annotations may sometimes differ from the golden standard, they are often rated more favorably by other humans. In contrast, some LLMs demonstrate closer alignment with the true labels but receive lower evaluations from experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。