arXiv:2509.15470cs.CVcs.AI2025-09被引 1

用无标签医学数据训练肺结节诊断模型,提升内部表现但外部泛化弱。

Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

  • 基于多模态影像与电子病历构建自监督预训练框架
  • 内部队列AUC达0.91,优于单模态与未正则化模型
  • 揭示模型在外部数据表现下降的潜在原因,适合医疗数据有限场景

肺结节诊断的多模态模型发展受限于标注数据稀缺及过拟合问题。本文利用纵向多模态医学档案开展自监督学习,构建联合嵌入预测架构(JEPA)进行预训练。在本机构内部队列中,微调后模型表现优于未正则化的多模态模型(AUC: 0.88)和仅影像模型(AUC: 0.73),达到0.91;但在外部队列中表现较差(本方法:AUC 0.72,仅影像模型:AUC 0.75)。研究构建了合成环境,分析JEPA性能下降的上下文因素。该方法利用无标签多模态医疗数据提升预测能力,同时揭示其优劣边界。

原文摘要 · Abstract (English)

The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-supervised learning from longitudinal and multimodal archives to address these challenges. We curate an unlabeled set of patients with CT scans and linked electronic health records from our home institution to power joint embedding predictive architecture (JEPA) pretraining. After supervised finetuning, we show that our approach outperforms an unregularized multimodal model and imaging-only model in an internal cohort (ours: 0.91, multimodal: 0.88, imaging-only: 0.73 AUC), but underperforms in an external cohort (ours: 0.72, imaging-only: 0.75 AUC). We develop a synthetic environment that characterizes the context in which JEPA may underperform. This work innovates an approach that leverages unlabeled multimodal medical archives to improve predictive models and demonstrates its advantages and limitations in pulmonary nodule diagnosis.

多模态学习自监督肺结节诊断医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。