arXiv:2411.15802eess.IVcs.AI2024-11被引 25

用2D自监督模型提升3D医学影像诊断准确率与可解释性

Medical Slice Transformer: Improved Diagnosis and Explainability on 3D Medical Images with DINOv2

  • 将DINOv2迁移到3D医学图像,通过切片融合实现跨维度分析
  • 在乳腺、胸部、膝关节三组数据上均优于3D ResNet,AUC最高达0.95
  • 生成的热力图更精准且符合解剖结构,适合临床医生信任使用

MRI和CT是诊断复杂疾病的重要横断面成像技术,但带标注的3D深度学习数据集稀缺。尽管如DINOv2等2D自监督方法表现优异,尚未应用于3D医学图像。此外,深度学习模型常因“黑箱”特性缺乏可解释性。本研究旨在将2D自监督模型(如DINOv2)拓展至3D医学图像分析,并评估其可解释潜力。提出Medical Slice Transformer(MST)框架,结合Transformer与2D特征提取器(DINOv2),用于3D医学图像处理。在三个临床数据集上对比3D残差网络(3D ResNet):乳腺MRI(651例)、胸部CT(722例)、膝关节MRI(1199例),分别用于乳腺癌诊断、肺结节良恶性预测及半月板撕裂检测。通过受试者工作特征曲线下面积(AUC)评估诊断性能,采用放射科医生对热力图的定性比较评估可解释性。结果表明,MST在所有数据集上AUC均高于ResNet:乳腺(0.94±0.01 vs. 0.91±0.02,P=0.02)、胸部(0.95±0.01 vs. 0.92±0.02,P=0.13)、膝关节(0.85±0.04 vs. 0.69±0.05,P=0.001)。MST生成的热力图在切片定位和病灶识别上更精确且解剖合理。

原文摘要 · Abstract (English)

MRI and CT are essential clinical cross-sectional imaging techniques for diagnosing complex conditions. However, large 3D datasets with annotations for deep learning are scarce. While methods like DINOv2 are encouraging for 2D image analysis, these methods have not been applied to 3D medical images. Furthermore, deep learning models often lack explainability due to their "black-box" nature. This study aims to extend 2D self-supervised models, specifically DINOv2, to 3D medical imaging while evaluating their potential for explainable outcomes. We introduce the Medical Slice Transformer (MST) framework to adapt 2D self-supervised models for 3D medical image analysis. MST combines a Transformer architecture with a 2D feature extractor, i.e., DINOv2. We evaluate its diagnostic performance against a 3D convolutional neural network (3D ResNet) across three clinical datasets: breast MRI (651 patients), chest CT (722 patients), and knee MRI (1199 patients). Both methods were tested for diagnosing breast cancer, predicting lung nodule dignity, and detecting meniscus tears. Diagnostic performance was assessed by calculating the Area Under the Receiver Operating Characteristic Curve (AUC). Explainability was evaluated through a radiologist's qualitative comparison of saliency maps based on slice and lesion correctness. P-values were calculated using Delong's test. MST achieved higher AUC values compared to ResNet across all three datasets: breast (0.94$\pm$0.01 vs. 0.91$\pm$0.02, P=0.02), chest (0.95$\pm$0.01 vs. 0.92$\pm$0.02, P=0.13), and knee (0.85$\pm$0.04 vs. 0.69$\pm$0.05, P=0.001). Saliency maps were consistently more precise and anatomically correct for MST than for ResNet. Self-supervised 2D models like DINOv2 can be effectively adapted for 3D medical imaging using MST, offering enhanced diagnostic accuracy and explainability compared to convolutional neural networks.

3D医学影像自监督学习可解释性DINOv2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。