对比卷积与注意力模型,提升电子显微镜下22类病毒识别准确率。
A Comparative and Hybrid Study of CNN and Transformer Models for Multi-Class Virus Classification in Transmission Electron Microscopy

- 用相同数据和训练条件比较CNN与Transformer模型性能。
- Swin Transformer达88.31%准确率,混合模型宏F1超85.28%。
- 模型差异互补可改善类别不平衡问题,适合医学图像分析者参考。
在透射电子显微镜(TEM)图像中自动识别病毒颗粒仍具挑战,主要因类别间相似性高、尺度变化大及严重类别不平衡。本研究使用TEM病毒数据集,对多种卷积神经网络与基于Transformer的架构进行对比评估,用于22类病毒分类。所有模型在相同预处理与优化条件下训练,并通过加权交叉熵缓解不平衡影响。性能以整体准确率及宏平均精确率、召回率、F1分数衡量。在独立模型中,Swin Transformer表现最优(准确率0.8831,宏F1 0.8444),其次为DeiT(准确率0.8669)。卷积模型表现相对较低,ResNet50在不平衡条件下准确率降至0.5887。为利用不同模型的表征优势,采用决策层融合策略。加权混合模型达到0.8831准确率与最高宏F1(0.8528),略优于均权混合配置。结果表明,模型架构异质性有助于提升类别间平衡性,且不牺牲整体预测精度。未来工作可探索尺度感知表示、特征级融合机制及扩展的TEM数据集,以进一步增强病毒识别的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
The automatic recognition of virus particles in transmission electron microscopy (TEM) images remains a demanding task, primarily owing to strong inter-class similarity, scale variability, and pronounced class imbalance. In this study, several convolutional neural networks and transformer-based architectures were comparatively evaluated for the classification of 22 virus categories using the TEM virus dataset. All models were trained under identical preprocessing and optimization conditions, and imbalance effects were mitigated through a weighted cross-entropy formulation. Performance was quantified using overall accuracy together with macro-averaged precision, recall, and F1 score. Among standalone models, the Swin Transformer achieved the highest accuracy (0.8831) and macro-F1 score (0.8444), followed by DeiT (accuracy 0.8669). Convolutional architectures exhibited comparatively lower balanced performance, with ResNet50 demonstrating substantial degradation (accuracy 0.5887) under imbalanced conditions. To exploit complementary representational properties, decision-level hybrid strategies were implemented. The performance-weighted hybrid attained an accuracy of 0.8831 and the highest macro-F1 score (0.8528), slightly surpassing the equal-weight hybrid configuration. These observations indicate that architectural heterogeneity contributes to improved inter-class balance without sacrificing overall predictive accuracy. Future work may explore scale-aware representations, feature-level fusion mechanisms, and expanded TEM datasets to further enhance robustness and generalization in virus identification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。