arXiv:2607.02569cs.CVcs.LG2026-07

用不确定性感知方法提升糖尿病视网膜病变筛查的可靠性

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

论文配图:Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift
图 1 · 摘自论文原文
  • 在RETFound模型基础上,通过贝叶斯化最后层头实现不确定性感知推理
  • 在APTOS数据集上,可使20%病例延迟诊断且假阴性归零,特异性仍高
  • 强调需跨数据集验证安全性能,仅靠准确率不足保障临床可信

本文针对糖尿病视网膜病变筛查,基于自监督视觉变换器模型RETFound(作为冻结特征编码器),在公开的APTOS 2019与DDR眼底图像数据集上,对多种不确定性感知的最后层适配方法进行安全性导向的实证评估。对比了缓存特征softmax头、事后温度校准、变分贝叶斯最后层头、对角拉普拉斯近似及SNGP风格缓存头。在APTOS数据集上,不确定性感知操作点显著提升敏感性与选择性转诊能力:约20%病例可被延迟处理,且已接受病例的假阴性降至零,同时保持高特异性。然而阈值调优也导致高假阳性代价下的假阴性下降,表明假阴性减少并非贝叶斯建模独有。在DDR数据集上,原生贝叶斯头虽呈现类似趋势但权衡较弱,而基于APTOS训练的SNGP检查点迁移效果差,未能提供有效的外部选择性转诊能力。结果凸显超越整体准确率的安全性评估重要性:不确定性感知最后层头可改善内部安全操作点,但可信的视网膜筛查结论必须依赖显式的安全覆盖评估与数据分布偏移下的第二数据集验证。

原文摘要 · Abstract (English)

This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using RETFound, a self-supervised vision-transformer retinal foundation model used here as a frozen feature encoder, and the public APTOS 2019 and DDR diabetic retinopathy fundus image datasets. We compare a cached-feature softmax head, post-hoc temperature scaling, variational Bayesian last-layer heads, a diagonal Laplace last-layer approximation, and an SNGP-style cached-feature head. On APTOS, uncertainty-aware operating points improved sensitivity and selective-referral behavior. The strongest APTOS selective-referral result deferred approximately 20 percent of cases and reduced accepted-case false negatives to zero while preserving high accepted-case specificity. However, threshold tuning also reduced false negatives at high false-positive cost, so false-negative reduction alone was not unique to Bayesian modeling. On DDR, native Bayesian heads qualitatively reproduced the APTOS direction but with weaker tradeoffs, while the APTOS-trained SNGP checkpoint transferred poorly and failed to provide useful external selective-referral behavior. These results highlight the value of safety-centered evaluation beyond aggregate accuracy: uncertainty-aware last-layer heads can improve internal safety-oriented operating points, but trustworthy retinal screening claims require explicit safety-coverage evaluation and second-dataset validation under shift.

糖尿病视网膜病变不确定性建模迁移学习医疗影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。