arXiv:2511.05529q-bio.QMcs.AI2025-11被引 2

用加权集成与不确定性判断,让糖尿病眼病筛查更准更可靠

Selective Diabetic Retinopathy Screening with Accuracy-Weighted Deep Ensembles and Entropy-Guided Abstention

  • 七种模型融合,按准确率加权投票,提升诊断鲁棒性
  • 过滤低置信度样本后准确率达99.44%,远超未过滤的93.70%
  • 适合需要高可靠性、可解释的医疗AI部署场景

糖尿病视网膜病变(DR)是糖尿病的微血管并发症,全球预计到2030年将影响超1.3亿人。早期识别对预防失明至关重要,但现有诊断依赖眼底照相和专家评审,成本高且资源紧张,加之其无症状特性,导致约25%病例被漏诊。尽管卷积神经网络(CNN)在医学影像中表现优异,但可解释性差且缺乏不确定性量化,影响临床可信度。本研究提出一种结合不确定性估计的深度集成学习框架,融合七种CNN架构——ResNet-50、DenseNet-121、MobileNetV3(Small/Large)及EfficientNet(B0/B2/B3),通过准确率加权多数投票融合输出。采用概率加权熵度量预测不确定性,使低置信度样本可被排除或标记复审。在35,000张EyePACS眼底图像上训练验证,未过滤准确率为93.70%(F1=0.9376);经不确定性过滤后,最高准确率达99.44%(F1=0.9932)。结果表明,基于不确定性的加权集成显著提升可靠性,同时保持高性能。该框架提供校准置信度输出与可调精度-覆盖权衡,为高风险医疗场景中的可信AI诊断提供通用范式。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR), a microvascular complication of diabetes and a leading cause of preventable blindness, is projected to affect more than 130 million individuals worldwide by 2030. Early identification is essential to reduce irreversible vision loss, yet current diagnostic workflows rely on methods such as fundus photography and expert review, which remain costly and resource-intensive. This, combined with DR's asymptomatic nature, results in its underdiagnosis rate of approximately 25 percent. Although convolutional neural networks (CNNs) have demonstrated strong performance in medical imaging tasks, limited interpretability and the absence of uncertainty quantification restrict clinical reliability. Therefore, in this study, a deep ensemble learning framework integrated with uncertainty estimation is introduced to improve robustness, transparency, and scalability in DR detection. The ensemble incorporates seven CNN architectures-ResNet-50, DenseNet-121, MobileNetV3 (Small and Large), and EfficientNet (B0, B2, B3)- whose outputs are fused through an accuracy-weighted majority voting strategy. A probability-weighted entropy metric quantifies prediction uncertainty, enabling low-confidence samples to be excluded or flagged for additional review. Training and validation on 35,000 EyePACS retinal fundus images produced an unfiltered accuracy of 93.70 percent (F1 = 0.9376). Uncertainty-filtering later was conducted to remove unconfident samples, resulting in maximum-accuracy of 99.44 percent (F1 = 0.9932). The framework shows that uncertainty-aware, accuracy-weighted ensembling improves reliability without hindering performance. With confidence-calibrated outputs and a tunable accuracy-coverage trade-off, it offers a generalizable paradigm for deploying trustworthy AI diagnostics in high-risk care.

糖尿病视网膜病变深度集成不确定性估计医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。