arXiv:2411.00832eess.IVcs.AI2024-11被引 9

混合CNN与视觉Transformer,实现骨肉瘤病理图像四分类高精度诊断

Advanced Hybrid Deep Learning Model for Enhanced Classification of Osteosarcoma Histopathology Images

  • 结合CNN局部特征提取与ViT全局模式捕捉,构建混合深度学习模型
  • 在TCIA数据集上达到99.08%准确率,首次实现四类精准分类
  • 为儿童骨肉瘤早期诊断提供可信赖的自动化工具,适合临床辅助决策

机器学习的进步正推动医学图像分析变革,尤其在癌症检测与分类方面。卷积神经网络(CNN)和视觉变压器(ViT)等技术现能精确分析复杂病理图像,实现自动检测并提升多种癌症的分类准确率。本研究聚焦于儿童和青少年最常见的骨癌——骨肉瘤(OS),其影响四肢长骨。早期且准确的检测对改善预后、降低死亡率至关重要。然而,癌症发病率上升及个性化治疗需求,使精准诊断与定制疗法面临挑战。本文提出一种新型混合模型,结合CNN与ViT,利用苏木精-伊红(H&E)染色病理图像提升骨肉瘤诊断准确率。CNN负责提取局部特征,ViT捕捉全局模式,二者融合后通过多层感知机(MLP)分类为四类:非肿瘤(NT)、非存活肿瘤(NVT)、存活肿瘤(VT)和非存活比例(NVR)。基于癌症影像档案(TCIA)数据集,模型取得99.08%准确率、99.10%精确率、99.28%召回率和99.23%F1分数。这是首次在该数据集上实现四分类,为骨肉瘤研究树立新基准,展现未来诊断进步的广阔前景。

原文摘要 · Abstract (English)

Recent advances in machine learning are transforming medical image analysis, particularly in cancer detection and classification. Techniques such as deep learning, especially convolutional neural networks (CNNs) and vision transformers (ViTs), are now enabling the precise analysis of complex histopathological images, automating detection, and enhancing classification accuracy across various cancer types. This study focuses on osteosarcoma (OS), the most common bone cancer in children and adolescents, which affects the long bones of the arms and legs. Early and accurate detection of OS is essential for improving patient outcomes and reducing mortality. However, the increasing prevalence of cancer and the demand for personalized treatments create challenges in achieving precise diagnoses and customized therapies. We propose a novel hybrid model that combines convolutional neural networks (CNN) and vision transformers (ViT) to improve diagnostic accuracy for OS using hematoxylin and eosin (H&E) stained histopathological images. The CNN model extracts local features, while the ViT captures global patterns from histopathological images. These features are combined and classified using a Multi-Layer Perceptron (MLP) into four categories: non-tumor (NT), non-viable tumor (NVT), viable tumor (VT), and none-viable ratio (NVR). Using the Cancer Imaging Archive (TCIA) dataset, the model achieved an accuracy of 99.08%, precision of 99.10%, recall of 99.28%, and an F1-score of 99.23%. This is the first successful four-class classification using this dataset, setting a new benchmark in OS research and offering promising potential for future diagnostic advancements.

病理图像深度学习骨肉瘤四分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。