用大模型自动分析社交媒体中的药物使用话题,准确率达90.9%
Automated Thematic Analyses Using LLMs: Xylazine Wound Management Social Media Chatter Use Case
- 将主题分析转为二分类任务,通过少量示例提示提升模型表现
- GPT-4o在验证集上准确率90.9%,高出现频主题与专家判断高度一致
- 适合需快速处理海量社交文本的公共卫生研究者使用
大型语言模型(LLMs)在归纳式主题分析中面临挑战,该任务需深度解读与领域专长。我们评估了使用LLMs复现专家主导的主题分析在社交媒体数据中的可行性。基于两个时间不重叠的Reddit数据集(分别为n=286和n=686),针对十二个由专家定义的主题,对比五种LLMs的表现。将任务建模为一系列二分类问题,采用零样本、单样本及少样本提示策略,以准确率、精确率、召回率和F1分数衡量性能。在验证集上,使用两样本提示的GPT-4o表现最佳(准确率:90.9%;F1分数:0.71)。对于高出现频率的主题,模型推导的主题分布与专家分类高度一致(如:xylazine使用:13.6% vs. 17.8%;MOUD使用:16.5% vs. 17.8%)。结果表明,少样本驱动的LLM方法可实现主题分析自动化,为定性研究提供可扩展的补充方案。
原文摘要 · Abstract (English)
Background Large language models (LLMs) face challenges in inductive thematic analysis, a task requiring deep interpretive and domain-specific expertise. We evaluated the feasibility of using LLMs to replicate expert-driven thematic analysis of social media data. Methods Using two temporally non-intersecting Reddit datasets on xylazine (n=286 and n=686, for model optimization and validation, respectively) with twelve expert-derived themes, we evaluated five LLMs against expert coding. We modeled the task as a series of binary classifications, rather than a single, multi-label classification, employing zero-, single-, and few-shot prompting strategies and measuring performance via accuracy, precision, recall, and F1-score. Results On the validation set, GPT-4o with two-shot prompting performed best (accuracy: 90.9%; F1-score: 0.71). For high-prevalence themes, model-derived thematic distributions closely mirrored expert classifications (e.g., xylazine use: 13.6% vs. 17.8%; MOUD use: 16.5% vs. 17.8%). Conclusions Our findings suggest that few-shot LLM-based approaches can automate thematic analyses, offering a scalable supplement for qualitative research. Keywords: thematic analysis, large language models, natural language processing, qualitative analysis, social media, prompt engineering, public health
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。