arXiv:2509.07795eess.IVcs.AI2025-09被引 2

改进的SegNet结合Grad-CAM,实现OCT眼底图像分层分割的高精度与可解释性。

Enhanced SegNet with Integrated Grad-CAM for Interpretable Retinal Layer Segmentation in OCT Images

  • 采用改进池化策略和混合损失函数提升对噪声图像的特征提取能力。
  • 在杜克OCT数据集上达到95.77%准确率,Dice系数0.9446,IoU达0.8951。
  • 集成Grad-CAM可视化,帮助医生理解模型决策,增强临床可信度。

光学相干断层扫描(OCT)对青光眼、糖尿病视网膜病变和老年黄斑变性等疾病的诊断至关重要。精准的视网膜分层分割可提供关键定量生物标志物,但人工分割耗时且变异大,传统深度学习模型常缺乏可解释性。本文提出一种基于SegNet的改进框架,实现自动化且可解释的视网膜分层分割。通过优化池化策略增强对噪声OCT图像的特征提取能力,采用融合分类交叉熵与Dice损失的混合损失函数,提升对薄层及不平衡结构的分割性能。引入梯度加权类激活映射(Grad-CAM)生成可视化解释,支持临床验证模型决策。在杜克OCT数据集上训练与验证,框架取得95.77%的验证准确率、0.9446的Dice系数和0.8951的交并比(IoU)。各层表现稳健,较薄边界仍具挑战。Grad-CAM可视化聚焦解剖相关区域,与临床生物标志物一致,显著提升透明性。该方法通过架构优化、定制损失与可解释AI的融合,实现了精度与可解释性的平衡,具备标准化OCT分析、提升诊断效率及增强临床对AI工具信任的潜力。

原文摘要 · Abstract (English)

Optical Coherence Tomography (OCT) is essential for diagnosing conditions such as glaucoma, diabetic retinopathy, and age-related macular degeneration. Accurate retinal layer segmentation enables quantitative biomarkers critical for clinical decision-making, but manual segmentation is time-consuming and variable, while conventional deep learning models often lack interpretability. This work proposes an improved SegNet-based deep learning framework for automated and interpretable retinal layer segmentation. Architectural innovations, including modified pooling strategies, enhance feature extraction from noisy OCT images, while a hybrid loss function combining categorical cross-entropy and Dice loss improves performance for thin and imbalanced retinal layers. Gradient-weighted Class Activation Mapping (Grad-CAM) is integrated to provide visual explanations, allowing clinical validation of model decisions. Trained and validated on the Duke OCT dataset, the framework achieved 95.77% validation accuracy, a Dice coefficient of 0.9446, and a Jaccard Index (IoU) of 0.8951. Class-wise results confirmed robust performance across most layers, with challenges remaining for thinner boundaries. Grad-CAM visualizations highlighted anatomically relevant regions, aligning segmentation with clinical biomarkers and improving transparency. By combining architectural improvements, a customized hybrid loss, and explainable AI, this study delivers a high-performing SegNet-based framework that bridges the gap between accuracy and interpretability. The approach offers strong potential for standardizing OCT analysis, enhancing diagnostic efficiency, and fostering clinical trust in AI-driven ophthalmic tools.

视网膜分割可解释AIOCT图像SegNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。