用医生描述引导模型预测肺结节恶性,更准更可信。
Vision-Language Model-Based Semantic-Guided Imaging Biomarker for Lung Nodule Malignancy Prediction
- 用医学语义引导预训练视觉语言模型,学习临床相关特征。
- 在NLST数据集上准确率(AUROC)达0.901,优于现有方法。
- 可解释性强,适用于不同医院数据,适合临床部署。
机器学习模型常依赖深度特征或语义特征评估肺结节恶性程度,但其推理时需人工标注、可解释性差且对影像变化敏感,限制了真实临床应用。本研究整合放射科医生对结节的语义描述,引导模型学习具有临床意义、鲁棒且可解释的影像特征以预测肺癌。使用来自国家肺癌筛查试验(NLST)的938例低剂量CT扫描(含1,261个结节)及语义特征,以及LIDC-IDRI数据集中的1,018例扫描(含2,625个病灶标注)。此外还获取了来自UCLA Health、LUNGx挑战赛和杜克大学肺癌筛查的三个外部数据集。采用参数高效微调方法对预训练对比语言-图像预训练(CLIP)模型进行微调,实现影像与语义文本特征对齐,并预测一年内肺癌诊断。该模型在NLST测试集中表现超越当前最优(SOTA)模型,AUROC为0.901,AUPRC为0.776;在外部数据集上也保持稳健。通过零样本推理,对结节边缘(AUROC: 0.807)、密度一致性(0.812)和胸膜附着(0.840)等语义特征亦取得良好预测性能。该方法在多中心数据中均优于现有模型,输出可解释,帮助医生理解模型决策依据,避免学习捷径,具备跨机构泛化能力。代码已开源:https://github.com/luotingzhuang/CLIP_nodule。
原文摘要 · Abstract (English)
Machine learning models have utilized semantic features, deep features, or both to assess lung nodule malignancy. However, their reliance on manual annotation during inference, limited interpretability, and sensitivity to imaging variations hinder their application in real-world clinical settings. Thus, this research aims to integrate semantic features derived from radiologists' assessments of nodules, guiding the model to learn clinically relevant, robust, and explainable imaging features for predicting lung cancer. We obtained 938 low-dose CT scans from the National Lung Screening Trial (NLST) with 1,261 nodules and semantic features. Additionally, the Lung Image Database Consortium dataset contains 1,018 CT scans, with 2,625 lesions annotated for nodule characteristics. Three external datasets were obtained from UCLA Health, the LUNGx Challenge, and the Duke Lung Cancer Screening. We fine-tuned a pretrained Contrastive Language-Image Pretraining (CLIP) model with a parameter-efficient fine-tuning approach to align imaging and semantic text features and predict the one-year lung cancer diagnosis. Our model outperformed state-of-the-art (SOTA) models in the NLST test set with an AUROC of 0.901 and AUPRC of 0.776. It also showed robust results in external datasets. Using CLIP, we also obtained predictions on semantic features through zero-shot inference, such as nodule margin (AUROC: 0.807), nodule consistency (0.812), and pleural attachment (0.840). Our approach surpasses the SOTA models in predicting lung cancer across datasets collected from diverse clinical settings, providing explainable outputs, aiding clinicians in comprehending the underlying meaning of model predictions. This approach also prevents the model from learning shortcuts and generalizes across clinical settings. The code is available at https://github.com/luotingzhuang/CLIP_nodule.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。