通过解耦双眼结构实现更精准的视网膜疾病诊断。
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis

- 用无监督方法将图像块分解为对应解剖区域的可解释槽。
- 在ODIR-5K上提升AUC 4.2点,显著优于基线模型。
- 适合关注可解释性与临床对比逻辑的医疗AI研究者。
视网膜诊断本质上具有双边性:医生常对比双眼同源结构(如视盘不对称)。然而多数深度学习模型仅处理单眼表示。本文探究显式结构对应是否有助于诊断,提出Anatomy-Slot方法。该方法通过无监督方式将图像块分解为一组涌现的、结构一致的槽,对应解剖区域,并利用双向交叉注意力对齐双眼槽。在ODIR-5K数据集上,使用n=10个随机种子,该方法相较匹配的ViT-L基线提升AUC 4.2点(95%置信区间;Wilcoxon符号秩检验,W=0,p=0.002)。通过破坏配对和高斯噪声测试,验证了对应关系的依赖性与鲁棒性。进一步报告了在REFUGE数据集上的视盘定位量化结果及交叉注意力定位分析。这些结果表明,以对象为中心的解剖对应关系为符合临床双边比较习惯的可解释诊断系统提供了可靠路径。
原文摘要 · Abstract (English)
Retinal diagnosis is inherently bilateral: clinicians compare homologous structures across eyes (e.g., optic disc asymmetry), yet most deep models operate on monocular representations. We investigate whether explicit structural correspondence improves diagnosis, and propose Anatomy-Slot to operationalize this hypothesis. Anatomy-Slot introduces an unsupervised anatomical bottleneck by decomposing patch tokens into a set of emergent, structurally-coherent slots that correspond to anatomical regions, then aligning these slots across eyes via bidirectional cross-attention. On ODIR-5K with $n=10$ seeds, the method improves AUC by $4.2$ points over a matched ViT-L baseline (95% CIs; Wilcoxon signed-rank test, $W=0$, $p=0.002$). Pairing disruption and stress testing under Gaussian noise provide controlled tests of correspondence dependence and robustness under corruption. We further report quantitative optic disc grounding on REFUGE and cross-attention localization analysis. Beyond the reported gains, these results indicate that object-centric anatomical correspondence offers a principled path toward interpretable diagnostic systems aligned with clinical bilateral comparison.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。