arXiv:2608.22417cs.AIcs.CL2026-08综述

GPT-5在问卷文本分析中表现接近人类,可作定性研究辅助工具。

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

  • 用提示工程让GPT-5进行归纳编码,与人类方法一致
  • 编码一致性ARI达0.61,主题生成0.54,接近人类内部一致性
  • 适合需要快速处理大量开放题的研究者使用

大型语言模型(LLMs)在质性研究文本分析中的应用日益增多,但其在归纳内容分析中的表现证据仍有限。本研究对比了人类与基于LLM(GPT-5.4)的归纳编码,分析了来自欧洲博士生调查的903份开放题回答,涵盖六个变量。五名人类编码员按标准方案进行归纳分析,而GPT-5.4采用既定提示流程执行相同任务。通过调整兰德指数(ARI)评估人类与LLM输出的一致性。结果显示,编码一致性的ARI为0.61,主题生成为0.54,接近人类内部一致性(ARI=0.68)和模型自身一致性(ARI=0.76)。不同变量间一致性差异显著,实体内一致性低者,实体间一致性也低,凸显数据特征与个体表现对可靠性的影响。总体表明,该设定下LLM可近似人类编码,尤其在编码层面,具备作为可扩展支持工具的潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.

文本分析大模型定性研究问卷分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。