用自监督方法提升角膜神经分割,助力糖尿病神经病变早期诊断
HMSViT: A Hierarchical Masked Self-Supervised Vision Transformer for Corneal Nerve Segmentation and Diabetic Neuropathy Diagnosis
- 分层掩码自监督视觉Transformer,多尺度提取特征且计算量低
- 在临床数据集上实现61.34%的神经分割mIoU和70.40%诊断准确率
- 减少对标注数据依赖,适合医疗图像小样本场景应用
糖尿病周围神经病变(DPN)影响近半数糖尿病患者,需早期发现。角膜共聚焦显微镜(CCM)可实现无创诊断,但现有自动方法存在特征提取效率低、依赖手工先验、数据不足等问题。本文提出HMSViT——一种分层掩码自监督视觉Transformer,用于角膜神经分割与DPN诊断。HMSViT采用基于池化的分层结构与双注意力机制,结合绝对位置编码,使浅层捕捉局部细节、深层融合全局上下文,同时降低计算成本。设计块掩码自监督学习框架,减少对标签数据依赖,提升特征鲁棒性;通过多尺度解码器融合层次特征,实现分割与分类。在临床CCM数据集上的实验表明,HMSViT达到61.34% mIoU(分割)和70.40%诊断准确率,优于Swin Transformer和HiViT等先进分层模型,性能提升最高达6.39%,且参数更少。消融实验证实,块掩码自监督与分层多尺度特征融合显著优于传统监督训练。综合实验验证了HMSViT在临床可落地性、鲁棒性与性能上的优势,具备实际部署潜力。
原文摘要 · Abstract (English)
Diabetic Peripheral Neuropathy (DPN) affects nearly half of diabetes patients, requiring early detection. Corneal Confocal Microscopy (CCM) enables non-invasive diagnosis, but automated methods suffer from inefficient feature extraction, reliance on handcrafted priors, and data limitations. We propose HMSViT, a novel Hierarchical Masked Self-Supervised Vision Transformer (HMSViT) designed for corneal nerve segmentation and DPN diagnosis. Unlike existing methods, HMSViT employs pooling-based hierarchical and dual attention mechanisms with absolute positional encoding, enabling efficient multi-scale feature extraction by capturing fine-grained local details in early layers and integrating global context in deeper layers, all at a lower computational cost. A block-masked self supervised learning framework is designed for the HMSViT that reduces reliance on labelled data, enhancing feature robustness, while a multi-scale decoder is used for segmentation and classification by fusing hierarchical features. Experiments on clinical CCM datasets showed HMSViT achieves state-of-the-art performance, with 61.34% mIoU for nerve segmentation and 70.40% diagnostic accuracy, outperforming leading hierarchical models like the Swin Transformer and HiViT by margins of up to 6.39% in segmentation accuracy while using fewer parameters. Detailed ablation studies further reveal that integrating block-masked SSL with hierarchical multi-scale feature extraction substantially enhances performance compared to conventional supervised training. Overall, these comprehensive experiments confirm that HMSViT delivers excellent, robust, and clinically viable results, demonstrating its potential for scalable deployment in real-world diagnostic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。