用大模型发现新冠后遗症三类临床亚型,助力个性化治疗
LLM-Augmented Computational Phenotyping of Long Covid
- 基于大模型迭代生成假设、提取证据、优化特征,挖掘长期数据
- 从1.35万人中识别出防护型、应答型、难治型三种亚型
- 适用于复杂慢性病研究,为精准医疗提供可复用框架
表型特征刻画对理解慢性病异质性及指导个体化干预至关重要。长期新冠是一种复杂且持续存在的病症,但其临床亚型仍不明确。本文提出一种名为「Grace Cycle」的大型语言模型增强计算表型框架,通过迭代整合假设生成、证据提取与特征优化,从纵向患者数据中发现具有临床意义的亚群。该框架在13,511名长期新冠患者中识别出三种显著不同的临床表型:保护型、应答型和难治型。这些表型在峰值症状严重程度、基线疾病负担以及纵向剂量反应模式上表现出明显差异,并在多个独立维度上获得强统计支持。本研究展示了大语言模型如何融入严谨、统计可靠的表型筛选流程,从复杂纵向数据中发现临床可解释的亚群。值得注意的是,该框架具备疾病无关性,可作为发现临床可解读亚型的一般方法。
原文摘要 · Abstract (English)
Phenotypic characterization is essential for understanding heterogeneity in chronic diseases and for guiding personalized interventions. Long COVID, a complex and persistent condition, yet its clinical subphenotypes remain poorly understood. In this work, we propose an LLM-augmented computational phenotyping framework ``Grace Cycle'' that iteratively integrates hypothesis generation, evidence extraction, and feature refinement to discover clinically meaningful subgroups from longitudinal patient data. The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory, based on 13,511 Long Covid participants. These phenotypes exhibit pronounced separation in peak symptom severity, baseline disease burden, and longitudinal dose-response patterns, with strong statistical support across multiple independent dimensions. This study illustrates how large language models can be integrated into a principled, statistically grounded pipeline for phenotypic screening from complex longitudinal data. Note that the proposed framework is disease-agnostic and offers a general approach for discovering clinically interpretable subphenotypes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。