arXiv:2509.09242cs.CVcs.AI2025-09被引 4

混合卷积与注意力机制,提升胃组织病理图像分类精度。

CoAtNeXt:An Attention-Enhanced ConvNeXtV2-Transformer Hybrid Model for Gastric Tissue Classification

  • 用ConvNeXtV2块替代原模型的MBConv层,增强特征提取能力。
  • 在两个数据集上达到96.47%~99.90%的高准确率和AUC值。
  • 适合医疗影像分析、病理诊断辅助系统研发人员参考。

胃病早期诊断对预防致命后果至关重要。尽管组织病理学检查仍是诊断金标准,但完全依赖人工,存在工作量大、不同病理医师间结果差异明显等问题,可能导致关键发现遗漏,且缺乏标准化流程影响一致性。为此,本研究提出一种新型混合模型CoAtNeXt,用于胃组织图像分类。该模型基于CoAtNet架构,将MBConv层替换为增强版ConvNeXtV2块,并引入卷积块注意力模块(CBAM),通过通道与空间注意力机制提升局部特征提取效果。模型经规模调整,在两个公开数据集上进行评估:HMU-GC-HE-30K(八分类)和GasHisSDB(二分类),并与10种CNN和10种ViT模型对比。结果显示,CoAtNeXt在HMU-GC-HE-30K上达96.47%准确率、96.60%精确率、96.47%召回率、96.45% F1值及99.89% AUC;在GasHisSDB上达98.29%准确率、98.07%精确率、98.41%召回率、98.23% F1值及99.90% AUC,优于所有对比模型并超越已有研究。实验表明,CoAtNeXt是一种鲁棒的胃组织病理分类架构,具备二分类与多分类能力,有潜力提升诊断准确性并减轻病理医师负担。

原文摘要 · Abstract (English)

Background and objective Early diagnosis of gastric diseases is crucial to prevent fatal outcomes. Although histopathologic examination remains the diagnostic gold standard, it is performed entirely manually, making evaluations labor-intensive and prone to variability among pathologists. Critical findings may be missed, and lack of standard procedures reduces consistency. These limitations highlight the need for automated, reliable, and efficient methods for gastric tissue analysis. Methods In this study, a novel hybrid model named CoAtNeXt was proposed for the classification of gastric tissue images. The model is built upon the CoAtNet architecture by replacing its MBConv layers with enhanced ConvNeXtV2 blocks. Additionally, the Convolutional Block Attention Module (CBAM) is integrated to improve local feature extraction through channel and spatial attention mechanisms. The architecture was scaled to achieve a balance between computational efficiency and classification performance. CoAtNeXt was evaluated on two publicly available datasets, HMU-GC-HE-30K for eight-class classification and GasHisSDB for binary classification, and was compared against 10 Convolutional Neural Networks (CNNs) and ten Vision Transformer (ViT) models. Results CoAtNeXt achieved 96.47% accuracy, 96.60% precision, 96.47% recall, 96.45% F1 score, and 99.89% AUC on HMU-GC-HE-30K. On GasHisSDB, it reached 98.29% accuracy, 98.07% precision, 98.41% recall, 98.23% F1 score, and 99.90% AUC. It outperformed all CNN and ViT models tested and surpassed previous studies in the literature. Conclusion Experimental results show that CoAtNeXt is a robust architecture for histopathological classification of gastric tissue images, providing performance on binary and multiclass. Its highlights its potential to assist pathologists by enhancing diagnostic accuracy and reducing workload.

病理图像注意力机制分类模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。