用AI分析癌症患者访谈,挖掘关键就医体验主题。
Analyzing Cancer Patients' Experiences with Embedding-based Topic Modeling and LLMs
- 用BERTopic和LLM结合提取患者故事中的主题
- 生物医学嵌入模型提升主题准确性和可解释性
- 适合医疗健康研究者与临床决策支持系统开发者
本研究利用神经主题建模与大语言模型分析13份癌症患者访谈(共132,722词),探索患者叙事中蕴含的深层主题。首先在相同预处理、分块与聚类配置下对比BERTopic与Top2Vec在关键词提取上的表现,随后使用GPT-4对单个访谈(I0)进行主题标注,并通过小规模人工评估(关注连贯性、清晰度、相关性)验证效果。结果表明BERTopic表现更优,故选用三种临床嵌入模型进一步实验。采用最佳设置(BioClinicalBERT)分析全部13份访谈,发现“癌症管理中的协调与沟通”和“治疗过程中的患者决策”是贯穿所有访谈的核心主题。尽管访谈为荷兰语机器翻译,且无临床专家参与评估,结果仍显示神经主题建模(尤其BERTopic)可为临床提供有价值的患者视角反馈,有助于提升文档导航效率并增强患者声音在医疗流程中的作用。
原文摘要 · Abstract (English)
This study investigates the use of neural topic modeling and LLMs to uncover meaningful themes from patient storytelling data, to offer insights that could contribute to more patient-oriented healthcare practices. We analyze a collection of transcribed interviews with cancer patients (132,722 words in 13 interviews). We first evaluate BERTopic and Top2Vec for individual interview summarization by using similar preprocessing, chunking, and clustering configurations to ensure a fair comparison on Keyword Extraction. LLMs (GPT4) are then used for the next step topic labeling. Their outputs for a single interview (I0) are rated through a small-scale human evaluation, focusing on {coherence}, {clarity}, and {relevance}. Based on the preliminary results and evaluation, BERTopic shows stronger performance and is selected for further experimentation using three {clinically oriented embedding} models. We then analyzed the full interview collection with the best model setting. Results show that domain-specific embeddings improved topic \textit{precision} and \textit{interpretability}, with BioClinicalBERT producing the most consistent results across transcripts. The global analysis of the full dataset of 13 interviews, using the BioClinicalBERT embedding model, reveals the most dominant topics throughout all 13 interviews, namely ``Coordination and Communication in Cancer Care Management" and ``Patient Decision-Making in Cancer Treatment Journey''. Although the interviews are machine translations from Dutch to English, and clinical professionals are not involved in this evaluation, the findings suggest that neural topic modeling, particularly BERTopic, can help provide useful feedback to clinicians from patient interviews. This pipeline could support more efficient document navigation and strengthen the role of patients' voices in healthcare workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。