arXiv:2608.16958eess.IVcs.CV2026-08中稿 · ECCT 2026

用混合模型提升低分辨率眼底图糖尿病分级准确率

ORViT-DR: Ordinally-Robust Hybrid ViT for Low-Resolution Diabetic Retinopathy Grading

论文配图:ORViT-DR: Ordinally-Robust Hybrid ViT for Low-Resolution Diabetic Retinopathy Grading
图 1 · 摘自论文原文
  • 结合卷积与Transformer,利用预训练骨干网络提取特征
  • 在28x28图像上达57%准确率,κ=0.5963,优于常规方法
  • 适合医疗影像中低分辨率、分级有序任务的场景

糖尿病视网膜病变(DR)是导致视力受损的主要原因之一。可靠的自动化分级系统可提高筛查的安全性与准确性。由于疾病进展具有渐进性,分级任务天然具有序数结构,相邻等级视觉特征相似。本文提出ORViT-DR,一种用于低分辨率眼底图像的混合深度学习框架。该方法通过预训练的ViT-Hybrid骨干网络,融合BiT-ResNetv2与Vision Transformer架构,实现局部特征提取与全局上下文建模。在MedMNISTv2数据集的RetinaMNIST子集(28×28眼底图像,五级严重程度标注)上验证,采用渐进式解冻、逐层学习率衰减、指数移动平均参数更新及集成预测策略。实验结果表明,该方法在官方测试集上达到57.00%分类准确率,二次加权κ系数为0.5963,宏平均F1得分为0.4293。结果说明,混合CNN-Transformer架构能有效捕捉序数眼底图像表征。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is one of the main causes of impaired vision. A good and reliable automated grading system can make the screening process safer and more accurate. Because DR stages progress gradually, the task of grading disease severity naturally follows an ordinal structure in which neighboring classes share similar visual characteristics. In this study, ORViT-DR, a hybrid deep learning framework, is designed to improve DR grading from low-resolution retinal images. The proposed approach combines convolutional feature extraction with transformer-based global context modeling through a pre-trained ViT-Hybrid backbone, which integrates BiT-ResNetv2 with a Vision Transformer architecture. The approach is tested on the RetinaMNIST subset of the MedMNISTv2 dataset, which contains 28x28 retinal fundus images annotated with five levels of disease severity. To promote stable training and better feature learning, the training strategy applies progressive layer unfreezing, layer-wise learning rate decay, exponential moving average (EMA) parameter updates, and ensemble-based prediction during inference. Experimental results on the official RetinaMNIST test set show that the proposed method achieves 57.00% classification accuracy, along with a quadratic weighted kappa score of 0.5963 and a macro-F1 score of 0.4293. These results suggest that hybrid CNN-Transformer architectures can provide effective representations for ordinal retinal image analysis.

糖尿病视网膜病变图像分级混合模型低分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。