用大模型分析儿童败血症数据,发现更精准的患者亚群。
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
- 将病历转为文本后用大模型生成嵌入,再聚类
- 最佳模型得分0.86,识别出营养与社会背景各异的亚群
- 适合资源有限地区做精准医疗决策参考
聚类患者亚群对个性化治疗和高效资源配置至关重要。传统方法难以处理高维异构医疗数据且缺乏上下文理解。本研究在低收入国家的儿科败血症数据集(2,686条记录,28个数值变量,119个分类变量)上评估基于大语言模型(LLM)的聚类方法,对比经典方法。患者记录被序列化为文本,含与不含聚类目标两种方式。使用量化版LLAMA 3.1 8B、DeepSeek-R1-Distill-Llama-8B(LoRA微调)及Stella-En-400M-V5模型生成嵌入,再以K均值聚类。经典方法包括基于UMAP的K-Medoids及FAMD降维后的混合数据聚类。通过轮廓系数与统计检验评估聚类质量与区分度。Stella-En-400M-V5达到最高轮廓系数0.86。含聚类目标的LLAMA 3.1 8B表现更优,能识别更多亚群,其特征涵盖营养、临床与社会经济维度。大模型方法优于传统技术,因能捕捉更丰富上下文并突出关键特征,表明其在资源受限环境下具备用于情境化表型分析与辅助决策的潜力。
原文摘要 · Abstract (English)
Clustering patient subgroups is essential for personalized care and efficient resource use. Traditional clustering methods struggle with high-dimensional, heterogeneous healthcare data and lack contextual understanding. This study evaluates Large Language Model (LLM) based clustering against classical methods using a pediatric sepsis dataset from a low-income country (LIC), containing 2,686 records with 28 numerical and 119 categorical variables. Patient records were serialized into text with and without a clustering objective. Embeddings were generated using quantized LLAMA 3.1 8B, DeepSeek-R1-Distill-Llama-8B with low-rank adaptation(LoRA), and Stella-En-400M-V5 models. K-means clustering was applied to these embeddings. Classical comparisons included K-Medoids clustering on UMAP and FAMD-reduced mixed data. Silhouette scores and statistical tests evaluated cluster quality and distinctiveness. Stella-En-400M-V5 achieved the highest Silhouette Score (0.86). LLAMA 3.1 8B with the clustering objective performed better with higher number of clusters, identifying subgroups with distinct nutritional, clinical, and socioeconomic profiles. LLM-based methods outperformed classical techniques by capturing richer context and prioritizing key features. These results highlight potential of LLMs for contextual phenotyping and informed decision-making in resource-limited settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。