用大模型做临床对话主题分析,需统一评估标准
Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models
- 提出三维度评估框架:有效性、可靠性、可解释性
- 发现现有研究评估方法差异大,难比较
- 适合医疗AI与质性研究交叉方向的学者参考
本文探讨大型语言模型(LLMs)在支持非结构化临床转录文本的主题分析中的应用。主题分析是挖掘患者与医护人员叙事模式的常用但资源密集的方法。我们系统回顾了近期将LLMs应用于主题分析的研究,并采访了一名执业临床医生。结果表明,当前方法在主题分析类型、数据集、提示策略和模型使用等方面均存在碎片化问题,尤其体现在评估环节。现有评估方法差异显著,从专家定性评审到自动相似性度量不等,阻碍了领域进展并妨碍跨研究基准比较。我们认为建立标准化评估实践对推动该领域至关重要。为此,我们提出一个以有效性、可靠性和可解释性为核心的评估框架。
原文摘要 · Abstract (English)
This position paper examines how large language models (LLMs) can support thematic analysis of unstructured clinical transcripts, a widely used but resource-intensive method for uncovering patterns in patient and provider narratives. We conducted a systematic review of recent studies applying LLMs to thematic analysis, complemented by an interview with a practicing clinician. Our findings reveal that current approaches remain fragmented across multiple dimensions including types of thematic analysis, datasets, prompting strategies and models used, most notably in evaluation. Existing evaluation methods vary widely (from qualitative expert review to automatic similarity metrics), hindering progress and preventing meaningful benchmarking across studies. We argue that establishing standardized evaluation practices is critical for advancing the field. To this end, we propose an evaluation framework centered on three dimensions: validity, reliability, and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。