通过异常驱动的自监督学习,提升放射影像表征效率与精度。
Abnormality-Driven Representation Learning for Radiology Imaging
- 利用病变增强对比学习,从CT切片中提取异常驱动的视觉表征
- 在肿瘤定位、肺病检测等任务上超越现有基础模型,且更省算力和数据
- 适合需要高效医学影像表征的临床研究与模型开发人员
目前放射影像深度学习主流采用在其他任务上预训练的3D网络,再针对具体任务微调。相比之下,病理学等邻近领域已成功应用基于自监督学习的任务无关基础模型,结合弱监督深度学习。但放射影像因3D成像的计算与数据需求高,以及解剖结构复杂,仍缺乏此类基础表征模型。为此,我们提出CLEAR框架,利用2D切片提取嵌入,并通过注意力聚合预测临床结局。其中引入病变增强对比学习(LeCL),以不同位置的CT切片中的异常为驱动,学习视觉表征。我们采用三种架构(ViT、VSSM、Gated CNN)进行单域对比学习。在肿瘤定位、肺病检测、患者分期三个临床任务上评估,对比四种前沿基础模型(包括BiomedCLIP)。结果表明,使用LeCL学习的表示,显著优于现有模型,且计算与数据消耗更低。
原文摘要 · Abstract (English)
To date, the most common approach for radiology deep learning pipelines is the use of end-to-end 3D networks based on models pre-trained on other tasks, followed by fine-tuning on the task at hand. In contrast, adjacent medical fields such as pathology, which focus on 2D images, have effectively adopted task-agnostic foundational models based on self-supervised learning (SSL), combined with weakly-supervised deep learning (DL). However, the field of radiology still lacks task-agnostic representation models due to the computational and data demands of 3D imaging and the anatomical complexity inherent to radiology scans. To address this gap, we propose CLEAR, a framework for radiology images that uses extracted embeddings from 2D slices along with attention-based aggregation for efficiently predicting clinical endpoints. As part of this framework, we introduce lesion-enhanced contrastive learning (LeCL), a novel approach to obtain visual representations driven by abnormalities in 2D axial slices across different locations of the CT scans. Specifically, we trained single-domain contrastive learning approaches using three different architectures: Vision Transformers, Vision State Space Models and Gated Convolutional Neural Networks. We evaluate our approach across three clinical tasks: tumor lesion location, lung disease detection, and patient staging, benchmarking against four state-of-the-art foundation models, including BiomedCLIP. Our findings demonstrate that CLEAR using representations learned through LeCL, outperforms existing foundation models, while being substantially more compute- and data-efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。