arXiv:2502.05517cs.CV2025-02被引 9

用Vision Transformer分析脑肺肾肿瘤影像,准确率达99.4%。

Evaluation of Vision Transformers for Multimodal Image Classification: A Case Study on Brain, Lung, and Kidney Tumors

  • 采用Swin Transformer和MaxViT模型,跨模态融合多类肿瘤影像数据。
  • 单数据集最高准确率99%,综合数据集达99.4%。
  • 适合医学影像分析、多模态模型研究者参考。

神经网络已成为癌症检测与分类的主流技术。本文评估了Vision Transformers(包括Swin Transformer和MaxViT)在磁共振成像(MRI)和计算机断层扫描(CT)影像数据集中的表现,涵盖脑瘤、肺部及肾脏肿瘤三类数据集,每类包含不同诊断标签,如脑胶质瘤、脑膜瘤、良恶性肺部病变以及囊肿、癌症等肾部异常。通过在单一与组合数据集上微调模型,实验显示Swin Transformer在各数据集上平均准确率达99%,联合数据集准确率达到99.4%。结果表明,基于Transformer的模型对多种影像模态与特征具有强适应性。但仍面临标注数据有限与可解释性差等挑战。未来将引入更多影像模态,提升诊断能力。跨数据集整合此类模型有望推动精准医疗发展,实现更高效全面的医疗解决方案。

原文摘要 · Abstract (English)

Neural networks have become the standard technique for medical diagnostics, especially in cancer detection and classification. This work evaluates the performance of Vision Transformers architectures, including Swin Transformer and MaxViT, in several datasets of magnetic resonance imaging (MRI) and computed tomography (CT) scans. We used three training sets of images with brain, lung, and kidney tumors. Each dataset includes different classification labels, from brain gliomas and meningiomas to benign and malignant lung conditions and kidney anomalies such as cysts and cancers. This work aims to analyze the behavior of the neural networks in each dataset and the benefits of combining different image modalities and tumor classes. We designed several experiments by fine-tuning the models on combined and individual datasets. The results revealed that the Swin Transformer provided high accuracy, achieving up to 99\% on average for individual datasets and 99.4\% accuracy for the combined dataset. This research highlights the adaptability of Transformer-based models to various image modalities and features. However, challenges persist, including limited annotated data and interpretability issues. Future work will expand this study by incorporating other image modalities and enhancing diagnostic capabilities. Integrating these models across diverse datasets could mark a significant advance in precision medicine, paving the way for more efficient and comprehensive healthcare solutions.

医学影像Vision Transformer多模态肿瘤分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。