arXiv:2604.17341cs.CVcs.AI2026-04

提出双分辨率注意力模型,量化评估糖尿病视网膜病变分级在跨域场景下的性能下降。

Dual-Resolution Attention-Gated Deep Learning with Ordinal Regression for Diabetic Retinopathy Grading: A Quantified Assessment of Cross-Domain Generalization

论文配图:Dual-Resolution Attention-Gated Deep Learning with Ordinal Regression for Diabetic Retinopathy Grading: A Quantified Assessment of Cross-Domain Generalization
图 1 · 摘自论文原文
  • 双分支网络分别处理不同分辨率和增强图像,用可学习门控融合特征。
  • 跨域测试中κ值下降0.202,93.7%预测仍与真实等级相差不超过一级。
  • 适合关注模型在真实筛查中泛化能力的研究者或临床部署人员。

糖尿病视网膜病变(DR)是导致可预防失明的主要原因,自动化分级可扩大筛查覆盖范围。然而,多数现有模型仅在训练数据集上验证,未能衡量其在真实筛查中因成像域差异带来的性能变化。本研究提出一种双分辨率分级框架,量化评估域偏移下的性能退化。两个EfficientNet主干网络处理同一眼底图像的不同视图:B0接收经Ben Graham标准化的224×224输入,强调血管结构;B3接收CLAHE增强的300×300输入,突出局灶性病变。一个可学习的注意力门对每张图像融合双分支输出,序数二元分解头将疾病严重程度建模为有序等级而非无序类别。训练使用4,149张图像(APTOS 2019,n=2,929;Messidor-2训练集,n=1,220);评估使用独立的APTOS测试集(n=733)和未参与训练及模型选择的Messidor-2测试集(n=524)。该运行在APTOS上达到0.882(95% CI 0.853–0.906)的加权二次κ,而在Messidor-2上为0.679(95% CI 0.613–0.735),显著下降0.202(95% CI 0.142–0.273);三次随机种子实验的平均κ为0.689 ± 0.021。关键发现:准确率下降19.3个百分点,但93.7%的预测仍与参考等级相差不超过一级;等级顺序在域偏移下依然保持,阈值定位则失效。可诊断性DR敏感度从0.879降至0.620。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a leading cause of preventable blindness, and automated grading could extend screening capacity. However, most reported DR models are validated only on the dataset they were trained on, leaving their behaviour under real screening variability unmeasured. This study presents a dual-resolution grading framework and quantifies how far performance falls when the imaging domain shifts. Two EfficientNet backbones process complementary views of each fundus image: B0 receives Ben Graham-normalised input at 224x224, emphasising vascular structure, while B3 receives CLAHE-enhanced input at 300x300, emphasising focal lesions. A learnable attention gate fuses the branches per image, and an ordinal binary-decomposition head models severity as an ordered scale rather than as unordered categories. Training used a combined set of 4,149 images (APTOS 2019, n = 2,929; Messidor-2 training portion, n = 1,220); evaluation used a held-out APTOS split (n = 733) and a Messidor-2 test set (n = 524) excluded from training and from all model selection. Quadratic weighted kappa was 0.882 (95% CI 0.853-0.906) on APTOS and 0.679 (95% CI 0.613-0.735) on Messidor-2 for this run, a significant gap of 0.202 (95% CI 0.142-0.273); across three random seeds the held-out kappa was 0.689 +/- 0.021. Critically, accuracy fell 19.3 points while 93.7% of predictions stayed within one grade of reference: ordering survives domain shift, threshold placement does not. Referable-DR sensitivity fell from 0.879 to 0.620.

眼科影像跨域泛化分级模型序数回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。