arXiv:2603.11403cs.CV2026-03

DeepHistoViT用视觉Transformer实现可解释的病理癌变分类,准确率达100%。

DeepHistoViT: An Interpretable Vision Transformer Framework for Histopathological Cancer Classification

  • 基于定制化视觉Transformer,融合注意力机制定位诊断关键区域。
  • 在肺癌、结肠癌数据集上准确率、F1等指标达100%,白血病数据集超99.8%。
  • 适合病理医生辅助诊断,结果可解释性强,临床实用价值高。

组织病理学是癌症诊断的金标准,因其能提供细胞级的组织形态评估。然而,人工阅片耗时耗力,且存在观察者间差异,亟需可靠的计算机辅助诊断工具。近年来,基于Transformer的深度学习模型在建模医学图像复杂空间依赖性方面展现出强大潜力。本文提出DeepHistoViT,一种用于自动分类组织病理图像的Transformer框架。该模型采用定制化视觉Transformer架构,并集成注意力机制,以捕捉细微细胞结构,同时通过注意力定位诊断相关区域提升可解释性。在三个公开的组织病理学数据集(涵盖肺癌、结肠癌和急性淋巴细胞白血病)上进行评估。实验结果显示,所有数据集均达到领先水平:肺癌与结肠癌数据集分类准确率、精确率、召回率、F1分数及ROC-AUC均为100%;急性淋巴细胞白血病数据集对应指标分别为99.85%、99.84%、99.86%、99.85%和99.99%,所有结果均附有95%置信区间。这些结果凸显了Transformer架构在组织病理图像分析中的有效性,并证明DeepHistoViT作为可解释的辅助诊断工具在临床决策支持中的潜力。

原文摘要 · Abstract (English)

Histopathology remains the gold standard for cancer diagnosis because it provides detailed cellular-level assessment of tissue morphology. However, manual histopathological examination is time-consuming, labour-intensive, and subject to inter-observer variability, creating a demand for reliable computer-assisted diagnostic tools. Recent advances in deep learning, particularly transformer-based architectures, have shown strong potential for modelling complex spatial dependencies in medical images. In this work, we propose DeepHistoViT, a transformer-based framework for automated classification of histopathological images. The model employs a customized Vision Transformer architecture with an integrated attention mechanism designed to capture fine-grained cellular structures while improving interpretability through attention-based localization of diagnostically relevant regions. The framework is evaluated on three publicly available histopathology datasets covering lung cancer, colon cancer, and acute lymphoblastic leukaemia. Experimental results demonstrate state-of-the-art performance across all datasets, with classification accuracy, precision, recall, F1-score, and ROC-AUC reaching 100 percent on the lung and colon cancer datasets, and 99.85 percent, 99.84 percent, 99.86 percent, 99.85 percent, and 99.99 percent respectively on the acute lymphoblastic leukaemia dataset. All performance metrics are reported with 95 percent confidence intervals. These results highlight the effectiveness of transformer-based architectures for histopathological image analysis and demonstrate the potential of DeepHistoViT as an interpretable computer-assisted diagnostic tool to support pathologists in clinical decision-making.

病理分析视觉Transformer可解释性癌症分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。