首个能理解肺功能检测图的多模态大模型,让AI诊断有理可依。
SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting
- 用专用编码器提取呼吸曲线特征,与数值数据对齐后输入大模型
- 诊断准确率AUROC达0.8977,在缺数据时仍100%出报告
- 适合需要可解释性医疗AI的临床研究与系统开发人员
慢性阻塞性肺病(COPD)是全球主要致残和致死的慢性呼吸系统疾病,其诊断依赖于肺功能测试(PFTs)中常规采集的呼吸波形时间序列。然而,当前多数AI模型仅输出分类结果而无诊断依据,且通用大语言模型尚无法理解呼吸波形,严重制约临床信任与应用。本文基于英国生物银行(UKB)234,028名个体数据,提出首个能理解呼吸波形的多模态大语言模型SpiroLLM。该模型通过SpiroEncoder提取呼吸曲线形态特征,并利用SpiroProjector将特征与肺功能数值对齐至统一潜在空间,最终驱动大模型生成完整诊断报告。实验表明,SpiroLLM诊断的AUROC为0.8977(95%置信区间:0.88–0.91)。在缺失核心数据的鲁棒性测试中,其有效响应率达100%,远超仅依赖文本的模型(13.4%),凸显多模态设计优势。本工作展示了生理信号与大语言模型深度融合的巨大潜力,为下一代可解释、可靠的临床决策支持工具树立新范式。
原文摘要 · Abstract (English)
Chronic Obstructive Pulmonary Disease (COPD), a major chronic respiratory disease with persistent airflow limitation, is a leading global cause of disability and mortality. Respiratory spirogram time series, routinely collected during pulmonary function tests (PFTs), play a critical role in the early detection of respiratory diseases and in monitoring lung function over time. However, most current AI models for COPD diagnosis are limited to outputting classification results without providing a rationale for their diagnostic process, while current Large Language Models (LLMs) cannot understand spirograms yet, which severely limits their clinical trust and adoption. To tackle this challenge, we leverage a cohort of 234,028 individuals from the UK Biobank (UKB) to propose SpiroLLM, the first multimodal large language model that can understand spirogram. The model extracts morphological features from respiratory curves via a SpiroEncoder and aligns them with PFT numerical values in a unified latent space using a SpiroProjector, ultimately empowering a large language model to generate a comprehensive diagnostic report. Experimental results confirm that SpiroLLM achieved a diagnostic AUROC of 0.8977 (95% CI: 0.88-0.91). In a robustness test with missing core data, it maintained a 100% valid response rate, far surpassing the 13.4% of a text-only model and showcasing the superiority of its multimodal design. This work demonstrates the substantial potential of deeply fusing physiological signals with large language models, establishing a new paradigm for the next generation of interpretable and reliable clinical decision support tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。