arXiv:2504.04025cs.CVcs.LG2025-04

对比视觉变压器与卷积网络在淋巴瘤诊断中的表现,两者均达100%准确率。

Artificial intelligence application in lymphoma diagnosis: from Convolutional Neural Network to Vision Transformer

  • 用视觉变压器和卷积网络分别处理淋巴瘤全切片图像进行分类
  • 在20例样本(各10例)上,两种模型测试准确率均为100%
  • 首次在相同数据集上直接比较二者性能,适合病理AI研究者参考

近期研究表明,当在足够大的数据集上预训练时,视觉变压器可超越卷积神经网络。视觉变压器在大规模数据集上表现优异,具备多模态训练特性。鉴于其出色的特征检测能力,我们探索了视觉变压器在间变性大细胞淋巴瘤与经典霍奇金淋巴瘤诊断中的应用,使用HE染色全切片图像。我们在同一数据集上对比了视觉变压器与先前设计的卷积神经网络的分类性能。数据集包含20例样本(每类10例),每例生成60个100×100像素、20倍放大下的图像块,共1200个图像块,其中90%用于训练,9%用于验证,10%用于测试。此前卷积神经网络模型的测试准确率为100%。本次视觉变压器模型的测试结果同样达到100%。据作者所知,这是首次在相同淋巴瘤数据集上直接比较视觉变压器与卷积神经网络的预测性能。总体而言,卷积神经网络架构更成熟,在缺乏大规模预训练时通常为首选。然而,本研究显示,即使在相对小规模的间变性大细胞淋巴瘤与经典霍奇金淋巴瘤数据集上,视觉变压器仍能实现与卷积神经网络相当且优异的准确率。

原文摘要 · Abstract (English)

Recently, vision transformers were shown to be capable of outperforming convolutional neural networks when pretrained on sufficiently large datasets. Vision transformer models show good accuracy on large scale datasets, with features of multi-modal training. Due to their promising feature detection, we aim to explore vision transformer models for diagnosis of anaplastic large cell lymphoma versus classical Hodgkin lymphoma using pathology whole slide images of HE slides. We compared the classification performance of the vision transformer to our previously designed convolutional neural network on the same dataset. The dataset includes whole slide images of HE slides for 20 cases, including 10 cases in each diagnostic category. From each whole slide image, 60 image patches having size of 100 by 100 pixels and at magnification of 20 were obtained to yield 1200 image patches, from which 90 percent were used for training, 9 percent for validation, and 10 percent for testing. The test results from the convolutional neural network model had previously shown an excellent diagnostic accuracy of 100 percent. The test results from the vision transformer model also showed a comparable accuracy at 100 percent. To the best of the authors' knowledge, this is the first direct comparison of predictive performance between a vision transformer model and a convolutional neural network model using the same dataset of lymphoma. Overall, convolutional neural network has a more mature architecture than vision transformer and is usually the best choice when large scale pretraining is not an available option. Nevertheless, our current study shows comparable and excellent accuracy of vision transformer compared to that of convolutional neural network even with a relatively small dataset of anaplastic large cell lymphoma and classical Hodgkin lymphoma.

病理诊断视觉变压器卷积网络淋巴瘤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。