arXiv:2504.07468eess.IVcs.CV2025-04被引 5

轻量级模型+新池化方法,提升肺炎与新冠胸片检测准确率

Novel Pooling-based VGG-Lite for Pneumonia and Covid-19 Detection from Imbalanced Chest X-Ray Datasets

  • 用轻量VGG-Lite结合边缘增强模块和2Max-Min池化,聚焦病灶边缘特征
  • 在不平衡数据集上达95%准确率、96.6% F1分数,优于主流模型
  • 适合医疗影像诊断场景,尤其对小样本疾病检测有实用价值

本文提出一种基于新型池化的轻量级VGG-Lite模型,以缓解胸片(CXR)数据集中的类别不平衡问题。深度学习自动检测肺炎自2020年新冠变异株出现以来成为研究热点,但标准卷积神经网络在医学数据中常受类别不平衡影响。所提模型创新包括:(I) 基于VGG-16与MobileNet-V2设计轻量级基础模型VGG-Lite;(II) 在其上引入并行分支的边缘增强模块(EEM),包含负片层及全新自定义池化层2Max-Min Pooling,该层专注于肺炎胸片中的边缘结构,作为高效空间注意力机制。实验在两个不同数据集上验证:首个来自公开网络,第二个由研究团队整合三个来源构建,更具挑战性。结果表明,所提框架在两个数据集上均显著优于预训练CNN模型及三种近期主流模型(视觉变压器、基于池化的视觉变压器、PneuNet)。在「肺炎不平衡胸片数据集」上,未使用任何预处理即达到95%准确率、97.1%精确率、96.1%召回率、96.6%F1分数。

原文摘要 · Abstract (English)

This paper proposes a novel pooling-based VGG-Lite model in order to mitigate class imbalance issues in Chest X-Ray (CXR) datasets. Automatic Pneumonia detection from CXR images by deep learning model has emerged as a prominent and dynamic area of research, since the inception of the new Covid-19 variant in 2020. However, the standard Convolutional Neural Network (CNN) models encounter challenges associated with class imbalance, a prevalent issue found in many medical datasets. The innovations introduced in the proposed model architecture include: (I) A very lightweight CNN model, `VGG-Lite', is proposed as a base model, inspired by VGG-16 and MobileNet-V2 architecture. (II) On top of this base model, we leverage an ``Edge Enhanced Module (EEM)" through a parallel branch, consisting of a ``negative image layer", and a novel custom pooling layer ``2Max-Min Pooling". This 2Max-Min Pooling layer is entirely novel in this investigation, providing more attention to edge components within pneumonia CXR images. Thus, it works as an efficient spatial attention module (SAM). We have implemented the proposed framework on two separate CXR datasets. The first dataset is obtained from a readily available source on the internet, and the second dataset is a more challenging CXR dataset, assembled by our research team from three different sources. Experimental results reveal that our proposed framework has outperformed pre-trained CNN models, and three recent trend existing models ``Vision Transformer", ``Pooling-based Vision Transformer (PiT)'' and ``PneuNet", by substantial margins on both datasets. The proposed framework VGG-Lite with EEM, has achieved a macro average of 95% accuracy, 97.1% precision, 96.1% recall, and 96.6% F1 score on the ``Pneumonia Imbalance CXR dataset", without employing any pre-processing technique.

肺炎检测胸片分析轻量模型图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。