arXiv:2410.00779cs.CVeess.IV2024-10

用自监督与知识蒸馏提升眼底病分级模型泛化能力

Local-to-Global Self-Supervised Representation Learning for Diabetic Retinopathy Grading

  • 结合自监督学习与知识蒸馏,设计局部到全局特征提取架构
  • 在眼底病数据集上实现79.1%线性分类准确率,超越同类模型
  • 无需删图、测试集超训练集50%,适合真实医疗场景部署

过去十年中,人工智能在图像分类与分割方面表现突出,但在真实临床数据上的表现仍逊于模拟数据。本研究提出一种新型混合学习模型,融合自监督学习与知识蒸馏,以增强模型的泛化性与鲁棒性。该模型采用ViT中的自注意力机制与标记(tokens),结合局部到全局的学习策略,从眼底图像中提取高维高质量特征空间。实验使用糖尿病视网膜病变数据集EyePACS,其结构复杂且病灶区域具有挑战性。本研究首次将自监督学习与知识蒸馏应用于该数据集。在模型设计上,首次实现测试集比训练集大50%的配置,且未移除任何图像。在多分类任务中,线性分类器准确率达79.1%,k-NN算法达74.36%,优于同类先进模型,并生成更优的表示空间。

原文摘要 · Abstract (English)

Artificial intelligence algorithms have demonstrated their image classification and segmentation ability in the past decade. However, artificial intelligence algorithms perform less for actual clinical data than those used for simulations. This research aims to present a novel hybrid learning model using self-supervised learning and knowledge distillation, which can achieve sufficient generalization and robustness. The self-attention mechanism and tokens employed in ViT, besides the local-to-global learning approach used in the hybrid model, enable the proposed algorithm to extract a high-dimensional and high-quality feature space from images. To demonstrate the proposed neural network's capability in classifying and extracting feature spaces from medical images, we use it on a dataset of Diabetic Retinopathy images, specifically the EyePACS dataset. This dataset is more complex structurally and challenging regarding damaged areas than other medical images. For the first time in this study, self-supervised learning and knowledge distillation are used to classify this dataset. In our algorithm, for the first time among all self-supervised learning and knowledge distillation models, the test dataset is 50% larger than the training dataset. Unlike many studies, we have not removed any images from the dataset. Finally, our algorithm achieved an accuracy of 79.1% in the linear classifier and 74.36% in the k-NN algorithm for multiclass classification. Compared to a similar state-of-the-art model, our results achieved higher accuracy and more effective representation spaces.

眼底病分级自监督学习知识蒸馏医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。