用可追溯的迭代方法自动分析临床访谈,提升主题分析的准确与可信度。
Automated Thematic Analysis for Clinical Qualitative Data: Iterative Codebook Refinement with Full Provenance
- 通过多轮迭代优化编码本,逐步提升主题一致性与复用性。
- 在五个数据集上四项指标优于六种基线方法,效果显著。
- 适合需要高可解释性与可复现性的临床研究团队使用。
主题分析(TA)广泛用于健康研究中从患者访谈中提取模式,但传统人工分析存在可扩展性差与可复现性低的问题。基于大语言模型的自动化方法虽有帮助,但现有方案生成的编码本泛化能力有限且缺乏分析过程审计能力。本文提出一种结合迭代编码本优化与全程溯源追踪的自动化主题分析框架。在涵盖临床访谈、社交媒体和公开转录文本的五个语料库上评估,该框架在四项数据集上达到最高综合质量评分,相比六种基线方法表现更优。多轮迭代使四个数据集的质量显著提升,效应量较大,主要得益于编码复用率与分布一致性增强,同时保持描述性质量。在两个儿科心脏病学临床语料中,自动生成的主题与专家标注高度一致。
原文摘要 · Abstract (English)
Thematic analysis (TA) is widely used in health research to extract patterns from patient interviews, yet manual TA faces challenges in scalability and reproducibility. LLM-based automation can help, but existing approaches produce codebooks with limited generalizability and lack analytic auditability. We present an automated TA framework combining iterative codebook refinement with full provenance tracking. Evaluated on five corpora spanning clinical interviews, social media, and public transcripts, the framework achieves the highest composite quality score on four of five datasets compared to six baselines. Iterative refinement yields statistically significant improvements on four datasets with large effect sizes, driven by gains in code reusability and distributional consistency while preserving descriptive quality. On two clinical corpora (pediatric cardiology), generated themes align with expert-annotated themes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。