arXiv:2507.08002cs.HCcs.AI2025-07被引 4

用大模型辅助心理研究编码,效率高但深度不足。

Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study

  • 用提示工程框架让大模型自动提取主题和编码。
  • 知识增强模型仅需10-15份访谈就达编码饱和,人类需90-99份。
  • 适合需快速分析大量文本的研究者,但需人工校验深度。

主题分析能深入理解参与者体验,但耗时耗力,限制其在大规模医疗研究中的应用。大语言模型(LLMs)可实现文本的规模化自动分析,有望解决此问题。本研究对比了基于GPT-4o的大模型与传统人工分析在医护人员减压试验访谈中的表现。采用RISEN提示工程框架,大模型与人工分析均完成编码、饱和点判断、片段标注(n=20)及主题归纳。结果表明:大模型生成的归纳性主代码与人工相似,但人工在归纳性子代码与主题整合上更优;知识增强型大模型在10-15份访谈内即达编码饱和,而基础模型和人工分别需15-20份与90-99份;基础模型识别的片段数量与人工相当,组间一致性达K=0.84,知识增强模型则产出较少片段。人工片段更长且含多重编码,大模型多为单一编码。总体而言,大模型显著提升分析效率,但缺乏人工深度。结合人类监督,大模型可有效推动心理健康研究的定性分析。

原文摘要 · Abstract (English)

Thematic analysis provides valuable insights into participants' experiences through coding and theme development, but its resource-intensive nature limits its use in large healthcare studies. Large language models (LLMs) can analyze text at scale and identify key content automatically, potentially addressing these challenges. However, their application in mental health interviews needs comparison with traditional human analysis. This study evaluates out-of-the-box and knowledge-base LLM-based thematic analysis against traditional methods using transcripts from a stress-reduction trial with healthcare workers. OpenAI's GPT-4o model was used along with the Role, Instructions, Steps, End-Goal, Narrowing (RISEN) prompt engineering framework and compared to human analysis in Dedoose. Each approach developed codes, noted saturation points, applied codes to excerpts for a subset of participants (n = 20), and synthesized data into themes. Outputs and performance metrics were compared directly. LLMs using the RISEN framework developed deductive parent codes similar to human codes, but humans excelled in inductive child code development and theme synthesis. Knowledge-based LLMs reached coding saturation with fewer transcripts (10-15) than the out-of-the-box model (15-20) and humans (90-99). The out-of-the-box LLM identified a comparable number of excerpts to human researchers, showing strong inter-rater reliability (K = 0.84), though the knowledge-based LLM produced fewer excerpts. Human excerpts were longer and involved multiple codes per excerpt, while LLMs typically applied one code. Overall, LLM-based thematic analysis proved more cost-effective but lacked the depth of human analysis. LLMs can transform qualitative analysis in mental healthcare and clinical research when combined with human oversight to balance participant perspectives and research resources.

主题分析大模型心理健康定性研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。