arXiv:2511.18627cs.CVcs.LG2025-11

用眼底图检测眼病,视觉变压器表现稳定,尤其适合早期病变识别。

Functional Localization Enforced Deep Anomaly Detection Using Fundus Images

  • 采用视觉变压器结合几何与色彩增强,提升跨数据集检测稳定性。
  • 糖尿病视网膜病变和老年黄斑变性检测准确率达78.9%~84.3%,青光眼仍易误判。
  • 生成对抗异常检测器可解释且泛化强,适合临床部署前的决策支持。

从眼底图像中可靠检测视网膜疾病面临成像质量差异、早期症状细微以及数据集间域偏移等挑战。本研究系统评估了在多种增强策略下,视觉变压器(ViT)在多个异构公开数据集及自建高质量眼底数据集AEyeDB上的表现。ViT在不同数据集和疾病上准确率范围为0.789至0.843,糖尿病视网膜病变和老年黄斑变性检测效果良好,而青光眼仍最易误判。几何与色彩增强带来最稳定的性能提升,直方图均衡化对结构细微特征主导的数据集有益,拉普拉斯增强则降低整体表现。在Papila数据集上,使用几何增强的ViT达到AUC 0.91,优于先前卷积集成基线(AUC 0.87),凸显了变压器架构与多数据集训练的优势。为补充分类器,开发基于GANomaly的异常检测器,实现AUC 0.76,具备重建可解释性及对未见数据的鲁棒泛化能力。通过GUESS进行概率校准,实现阈值无关的决策支持,便于未来临床应用。

原文摘要 · Abstract (English)

Reliable detection of retinal diseases from fundus images is challenged by the variability in imaging quality, subtle early-stage manifestations, and domain shift across datasets. In this study, we systematically evaluated a Vision Transformer (ViT) classifier under multiple augmentation and enhancement strategies across several heterogeneous public datasets, as well as the AEyeDB dataset, a high-quality fundus dataset created in-house and made available for the research community. The ViT demonstrated consistently strong performance, with accuracies ranging from 0.789 to 0.843 across datasets and diseases. Diabetic retinopathy and age-related macular degeneration were detected reliably, whereas glaucoma remained the most frequently misclassified disease. Geometric and color augmentations provided the most stable improvements, while histogram equalization benefited datasets dominated by structural subtlety. Laplacian enhancement reduced performance across different settings. On the Papila dataset, the ViT with geometric augmentation achieved an AUC of 0.91, outperforming previously reported convolutional ensemble baselines (AUC of 0.87), underscoring the advantages of transformer architectures and multi-dataset training. To complement the classifier, we developed a GANomaly-based anomaly detector, achieving an AUC of 0.76 while providing inherent reconstruction-based explainability and robust generalization to unseen data. Probabilistic calibration using GUESS enabled threshold-independent decision support for future clinical implementation.

眼底图像异常检测视觉变压器医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。