针对胃肠影像复杂异常,提出高效实时分析方法
Abnormalities and Disease Detection in Gastro-Intestinal Tract Images
- 融合纹理与局部二值模式设计轻量网络,兼顾速度与精度
- 在HyperKvasir数据集上达41 FPS、F1=0.91、准确率0.99
- 适合临床实时诊断场景,尤其适用于低算力设备
胃肠道(GI)图像分析在医学诊断中至关重要。本研究针对传统方法因异常多样性与复杂性难以实现精准分类与分割的问题,提出适用于实时应用的高效解决方案。初期采用基于纹理的特征提取与分类方法,在Kvasir V2数据集上实现超过4000 FPS的处理速度,达到F1分数0.76和准确率0.98。随后转向深度学习,结合数据袋装技术优化模型,在HyperKvasir数据集上获得0.92准确率与0.60 F1分数,于Kvasir V2上达0.88 F1分数。为支持实时检测,设计集成纹理与局部二值模式的轻量级神经网络,通过学习阈值缓解类间相似与类内差异,实现在HyperKvasir上41 FPS、0.99准确率与0.91 F1分数。此外,提出两种分割工具,利用深度可分离卷积与神经网络集成,提升低帧率场景下的检测性能。整体研究从传统纹理方法过渡至深度学习与集成策略,构建了可扩展的胃肠图像分析框架。
原文摘要 · Abstract (English)
Gastrointestinal (GI) tract image analysis plays a crucial role in medical diagnosis. This research addresses the challenge of accurately classifying and segmenting GI images for real-time applications, where traditional methods often struggle due to the diversity and complexity of abnormalities. The high computational demands of this domain require efficient and adaptable solutions. This PhD thesis presents a multifaceted approach to GI image analysis. Initially, texture-based feature extraction and classification methods were explored, achieving high processing speed (over 4000 FPS) and strong performance (F1-score: 0.76, Accuracy: 0.98) on the Kvasir V2 dataset. The study then transitions to deep learning, where an optimized model combined with data bagging techniques improved performance, reaching an accuracy of 0.92 and an F1-score of 0.60 on the HyperKvasir dataset, and an F1-score of 0.88 on Kvasir V2. To support real-time detection, a streamlined neural network integrating texture and local binary patterns was developed. By addressing inter-class similarity and intra-class variation through a learned threshold, the system achieved 41 FPS with high accuracy (0.99) and an F1-score of 0.91 on HyperKvasir. Additionally, two segmentation tools are proposed to enhance usability, leveraging Depth-Wise Separable Convolution and neural network ensembles for improved detection, particularly in low-FPS scenarios. Overall, this research introduces novel and adaptable methodologies, progressing from traditional texture-based techniques to deep learning and ensemble approaches, providing a comprehensive framework for advancing GI image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。