用注意力与小波变换提升胃镜图像异常分类准确率
Capsule Endoscopy Multi-classification via Gated Attention and Wavelet Transformations
- 结合门控注意力与小波变换,聚焦图像关键区域,抑制噪声
- 在不平衡数据集上达94.81%平衡准确率,F1-score达91.11%
- 适合医疗影像分析、内窥镜自动诊断场景的从业者参考
胃肠道异常显著影响患者健康,需及时诊断以有效治疗。为此,从胶囊内镜(VCE)视频帧中实现异常自动分类对优化诊断流程至关重要。本文提出一种新模型,通过将全维门控注意力(OGA)机制与小波变换技术融入架构,使模型能聚焦于内镜图像中的关键区域,降低噪声与无关特征干扰。该方法在纹理与色彩高度多变的胶囊内镜图像中尤为有效。小波变换高效捕捉空间与频域信息,增强特征提取能力,尤其利于发现微小异常。同时,将平稳小波变换(SWT)与离散小波变换(DWT)提取的特征进行通道拼接,以捕获多尺度特征,对检测息肉、溃疡和出血等病变至关重要。该方法提升了在不平衡数据集上的分类性能。模型训练与验证准确率分别为92.76%与91.19%,损失为0.2057与0.2700。平衡准确率94.81%,AUC 87.49%,F1-score 91.11%,精确率91.17%,召回率91.19%,特异性98.44%。模型性能优于VGG16与ResNet50基线模型,证明其在识别多种胃肠道异常方面的优越性。
原文摘要 · Abstract (English)
Abnormalities in the gastrointestinal tract significantly influence the patient's health and require a timely diagnosis for effective treatment. With such consideration, an effective automatic classification of these abnormalities from a video capsule endoscopy (VCE) frame is crucial for improvement in diagnostic workflows. The work presents the process of developing and evaluating a novel model designed to classify gastrointestinal anomalies from a VCE video frame. Integration of Omni Dimensional Gated Attention (OGA) mechanism and Wavelet transformation techniques into the model's architecture allowed the model to focus on the most critical areas in the endoscopy images, reducing noise and irrelevant features. This is particularly advantageous in capsule endoscopy, where images often contain a high degree of variability in texture and color. Wavelet transformations contributed by efficiently capturing spatial and frequency-domain information, improving feature extraction, especially for detecting subtle features from the VCE frames. Furthermore, the features extracted from the Stationary Wavelet Transform and Discrete Wavelet Transform are concatenated channel-wise to capture multiscale features, which are essential for detecting polyps, ulcerations, and bleeding. This approach improves classification accuracy on imbalanced capsule endoscopy datasets. The proposed model achieved 92.76% and 91.19% as training and validation accuracies respectively. At the same time, Training and Validation losses are 0.2057 and 0.2700. The proposed model achieved a Balanced Accuracy of 94.81%, AUC of 87.49%, F1-score of 91.11%, precision of 91.17%, recall of 91.19% and specificity of 98.44%. Additionally, the model's performance is benchmarked against two base models, VGG16 and ResNet50, demonstrating its enhanced ability to identify and classify a range of gastrointestinal abnormalities accurately.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。