用高效模型提升胃镜图像分类准确率,还让结果可解释。
Enhanced Multi-Class Classification of Gastrointestinal Endoscopic Images with Interpretable Deep Learning Model
- 基于EfficientNetB3构建轻量模型,不依赖数据增强。
- 在Kvasir数据集上达到94.25%测试准确率。
- 结合LIME可视化关键病灶区域,适合医疗AI研究者。
内镜检查是评估胃肠道的重要手段,在识别胃肠道疾病中起关键作用。近年来深度学习在复杂模型与数据增强方法推动下显著提升了异常检测能力。本研究基于Kvasir数据集中的8000张标注胃镜图像,涵盖8个类别,提出一种新方法以提升分类精度。采用EfficientNetB3作为主干网络,在不依赖数据增强的前提下保持适中模型复杂度,测试准确率达94.25%,精确率和召回率分别为94.29%和94.24%。此外,通过局部可解释模型无关解释(LIME)生成显著性图,定位影响预测的关键图像区域,增强模型可解释性。研究表明,该工作将高精度与可解释性结合,推动AI在医学影像中的应用。
原文摘要 · Abstract (English)
Endoscopy serves as an essential procedure for evaluating the gastrointestinal (GI) tract and plays a pivotal role in identifying GI-related disorders. Recent advancements in deep learning have demonstrated substantial progress in detecting abnormalities through intricate models and data augmentation methods.This research introduces a novel approach to enhance classification accuracy using 8,000 labeled endoscopic images from the Kvasir dataset, categorized into eight distinct classes. Leveraging EfficientNetB3 as the backbone, the proposed architecture eliminates reliance on data augmentation while preserving moderate model complexity. The model achieves a test accuracy of 94.25%, alongside precision and recall of 94.29% and 94.24% respectively. Furthermore, Local Interpretable Model-agnostic Explanation (LIME) saliency maps are employed to enhance interpretability by defining critical regions in the images that influenced model predictions. Overall, this work highlights the importance of AI in advancing medical imaging by combining high classification accuracy with interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。